Researchers release CLIR-Bench to evaluate multimodal QA over irregular clinical time series
Benchmark comprises 6,600 QA instances from de-identified ICU records and targets models’ ability to ground answers in sparse, irregular temporal evidence.
1 source · cross-referenced
- Introduces CLIR-Bench, a benchmark for multimodal question answering over irregular clinical time series.
- Dataset contains 6,600 QA instances spanning 11 clinical variables, organized into four capability dimensions and 11 tasks.
- Each question is linked to explicit temporal evidence and task-specific answer derivation rules.
- Experiments indicate existing generalist models struggle to retrieve and reason over sparse clinical evidence.
Researchers from multiple institutions introduced CLIR-Bench, a benchmark designed to evaluate multimodal question answering over irregular clinical time series. The benchmark is constructed from de-identified ICU records using a four-stage pipeline, emphasizing the challenges posed by sparse, irregularly sampled, and asynchronous clinical data.
CLIR-Bench comprises 6,600 QA instances spanning 11 clinical variables and is organized into four capability dimensions and 11 tasks. Each question is explicitly linked to temporal evidence and task-specific answer derivation rules, enabling evaluation of both answer correctness and evidence grounding.
Preliminary experiments reported in the paper indicate that existing generalist models struggle to retrieve and reason over sparse clinical evidence, underscoring the need for stronger methods tailored to irregular time-series reasoning.
The authors provide the dataset and code via Hugging Face, facilitating reproducibility and further research into clinical multimodal QA systems.
- Aug 28, 2026 · Google DeepMind — Blog
Google DeepMind pilots first double-blind AI evaluations with cryptographic safeguards
Trust79 - Aug 26, 2026 · arXiv cs.AI
New NL2SQL benchmark shows enterprise database complexity degrades model performance
Trust79 - Aug 26, 2026 · arXiv cs.AI
New benchmark shows how reader-facing memory formats affect LLM evaluation scores
Trust79