ClinLens benchmark exposes gap between executable and correct clinical coding agents
A new benchmark of 200 executable clinical tasks shows top standardized model-scaffold achieves 56.3% scope-macro STRICTPASS, while five biomedical systems adapted to GPT-4o-mini reach at most 2.9%.
1 source · cross-referenced
- ClinLens introduces 200 executable tasks across five linked MIMIC resources, spanning structured EHRs, notes, ECGs, chest radiographs, and echocardiograms.
- The strongest of 24 standardized model-scaffold configurations achieves 56.3% scope-macro STRICTPASS on a fixed 126-task suite despite 100% EXECSUCCESS.
- Five biomedical systems adapted to GPT-4o-mini reach at most 2.9% scope-macro STRICTPASS on the same suite.
- A separately configured coding agent solves 83 of 126 tasks, highlighting a gap between runnable submissions and correct clinical analyses.
Researchers introduced ClinLens, a benchmark designed to evaluate clinical data-science agents on heterogeneous longitudinal records, including structured electronic health records, clinical notes, electrocardiograms, chest radiographs, and echocardiograms drawn from five linked MIMIC resources.
The benchmark comprises 200 executable tasks organized under a 4 x 5 taxonomy that crosses four patient-time scopes with five analysis capabilities, enabling standardized evaluation of agents' ability to handle multimodal and temporal clinical data.
Evaluation uses a program-first reverse synthesis approach, where each task is paired with an evaluator-private reference workflow and checked for required artifacts, cohort and temporal semantics, and the final answer, ensuring auditable and reproducible analyses.
On a fixed 126-task suite, the strongest of 24 standardized model-scaffold configurations achieved 56.3% scope-macro STRICTPASS while maintaining 100% EXECSUCCESS, indicating that while all submissions executed without error, only about half met correctness criteria across patient-time scopes.
By comparison, a separately configured coding agent solved 83 of the 126 tasks, while five biomedical systems adapted to GPT-4o-mini reached at most 2.9% scope-macro STRICTPASS, underscoring a substantial performance gap between general-purpose coding agents and domain-adapted biomedical systems.
The authors note that these results expose a meaningful divide between runnable submissions and correct clinical analyses, emphasizing the need for benchmarks that stress both executability and clinical validity in longitudinal multimodal settings.
- Jul 31, 2026 · arXiv cs.AI
Study finds objective misalignment undermines LLM multi-agent systems in adversarial settings
Trust79 - Jul 31, 2026 · arXiv cs.AI
RL fine-tuning yields more structured internal representations than SFT for mathematical reasoning, study finds
Trust79 - Jul 30, 2026 · Google DeepMind — Blog
Google DeepMind launches Gemini Robotics ER 2 for real-time robot orchestration and multi-robot collaboration
Trust79