Researchers release CLIR-Bench to evaluate multimodal QA over irregular clinical time series
Introduces CLIR-Bench, a benchmark for multimodal question answering over irregular clinical time series.
Introduces CLIR-Bench, a benchmark for multimodal question answering over irregular clinical time series.
An applied AI engineer compared 100 human-annotated traces with automated eval systems to assess their reliability.
OpenAI published an analysis identifying reliability and accuracy concerns in SWE-Bench Pro, a popular benchmark for evaluating AI coding performance.
AgentLens introduces a benchmark that evaluates interactive coding agents by their full task trajectory rather than a binary pass/fail outcome.
A new benchmark called CSTutorBench evaluates how well small language models can act as tutors in K-12 computer science education using block-based programming in VEX VR.
ScarfBench introduces 34 enterprise Java applications, 204 migration tasks, and 1,331 expert-written tests to evaluate AI agents on framework modernization.
GPTNT is a new benchmark for evaluating real-time collaboration between multimodal agents using the cooperative video game 'Keep Talking and Nobody Explodes'.
Accuracy saturation in AI benchmarks often leads to retirement or replacement, but this misses other performance dimensions like construct validity, efficiency, and reliability.
A new benchmark introduces a contamination-aware, multi-zone protocol to evaluate when LLMs should answer or abstain.
A nine-judge LLM-as-a-judge panel provides only about two independent votes’ worth of information due to correlated errors.