WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling
This arXiv update presents WMLLM, a self-evolving optimization-agent framework that uses predict-then-act world modeling. The approach leverages large language models to forecast promising optimization directions before candidate generation, followed by agentic refinement, population-based search, and reinforcement learning to improve both the world model and optimization strategy. Experiments on black-box optimization, especially multi-objective molecular optimization, demonstrate improved sample efficiency and competitive final performance, achieving state-of-the-art results under limited evaluation budgets.
Why it matters: If the findings generalize, WMLLM could reduce evaluation costs in complex optimization problems. This suggests a path to more efficient design cycles in domains relying on expensive experiments or simulations.
AI-assisted brief
official source