DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation
DS-Lighting introduces an explicit, four-layer harness (data, workflow, execution, evaluation) to standardize how data-science tasks are represented and executed by LLM agents. This design supports both predefined pipelines and adaptive search, aiming to improve end-to-end robustness.
By integrating multiple open-source benchmarks into a shared MLE-Bench-like task format, the framework enables controlled comparisons under a common interface, sandboxed runtime, and metric protocol. The result is improved reproducibility, comparability, and reliability across heterogeneous tasks and models.
The work includes empirical evaluations across agents, harness configurations, models, and ablations, and provides a public codebase at the linked repository for reproducibility and adoption.
AI-assisted brief
source-only
Add a comment
No account required