Google DeepMind pilots first double-blind AI evaluations with cryptographic safeguards
Google DeepMind launched the first double-blind evaluation of a proprietary frontier AI model using cryptographically secure environments.
Google DeepMind launched the first double-blind evaluation of a proprietary frontier AI model using cryptographically secure environments.
Introduces ESQ-Bench, an Oracle-first NL2SQL benchmark with three enterprise schema complexity tiers and silent-divergence evaluation.
A new benchmark called RENDER evaluates how the format of memory inputs (raw dialogue, summaries, structured records) affects LLM performance scores.
Wazobia Eval is the first benchmark focused on Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning.
Eleven widely used open-source ASR models reproduced benchmark errors and omitted audible words when acoustic cues suggested the test context.
Backtrader-Bench is a new benchmarking framework for evaluating LLM coding agents in algorithmic trading using self-generated multiple-choice questions (MCQs).
Apple’s ML Research team introduced DeepAmbigQA, a dataset of 3,600 questions designed to test LLMs on ambiguous multi-hop reasoning.
smevals is a new open-source eval suite for testing AI models, prompts, and agent harnesses.
A new arXiv paper introduces a consensus-based evaluation framework that ranks LLM responses by aggregating peer preferences rather than relying on fixed ground truth.
Apple’s Machine Learning Research team introduced LVSum, a human-annotated benchmark for evaluating long-form video summarization with fine-grained temporal alignment.