PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems
The paper presents a six-dimensional scoring framework across four difficulty regimes and multiple scene modalities, including standard, object creation, and hidden objects, plus a manipulation category.
Across seven models, reported accuracies range from 25% to 67%, with common failures attributed to premature responses, inefficient exploration, and inconsistent grounding in simulator feedback rather than fundamental understanding.
Why it matters: Demonstrates the gap between static reasoning and interactive physical reasoning; highlights the need for better multi-step tool use and grounding to improve LLM performance in dynamic environments.
AI-assisted brief
official source
Add a comment
No account required