A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making
The Oncology Decision Boundary Benchmark (ODBB) evaluates 2,005 decision points drawn from NCCN guidelines and colorectal cancer cases across nine frontier LLMs released between 2025 and 2026. A deterministic scorer classified 14 failure types, with oncologist validation showing substantial decision-path blind spots when choosing guideline pathways before detailed reasoning.
Key findings indicate that a sizable portion of items are not answered correctly by any model, and some models give unsafe commitments or fail to commit to a correct step despite knowing it. The results argue that model quality alone is insufficient for clinical deployment; effective decision support will require architectures that can detect competence boundaries and route decisions to clinicians.
AI-assisted brief
source-only
Add a comment
No account required