WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling
WMLLM integrates world modeling with a predict-then-act paradigm to guide candidate generation in black-box optimization tasks. The agent predicts promising directions and then generates candidates, with multi-turn refinements to continually improve its model and strategy.
The method emphasizes sample efficiency and leverages population-based search and reinforcement learning to evolve its optimization approach. Empirical results are highlighted in multi-objective molecular optimization benchmarks under constrained budgets.
Why it matters: If the findings generalize, WMLLM could reduce evaluation costs in complex optimization problems. This suggests a path to more efficient design cycles in domains relying on expensive experiments or simulations.
AI-assisted brief
official source
Add a comment
No account required