Google DeepMind pilots first double-blind AI evaluations with cryptographic safeguards
A pilot program with the Singapore AI Safety Institute and partners tests a proprietary Gemini model against confidential benchmarks in a privacy-preserving environment to prevent benchmark contamination.
1 source · cross-referenced
- Google DeepMind launched the first double-blind evaluation of a proprietary frontier AI model using cryptographically secure environments.
- The pilot tests a Gemini Flash Lite model against confidential benchmarks in collaboration with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons.
- The approach uses Confidential Space in Google Cloud to ensure neither the model weights nor the evaluation prompts are visible to the other party.
- The goal is to prevent benchmark contamination and increase trust in AI evaluations for high-stakes use cases like cybersecurity or government assessments.
Google DeepMind announced a pilot program to run the world’s first double-blind evaluation of a proprietary frontier AI model, using cryptographically secure environments to prevent benchmark contamination. The effort tests a Gemini Flash Lite model against confidential benchmarks in a privacy-preserving setup, where neither the model provider nor the evaluator can access the other’s sensitive data.
The pilot is a collaboration with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. By leveraging Confidential Space within Google Cloud’s Confidential Computing portfolio, the initiative cryptographically verifies that the external evaluation data and the proprietary model remain private to their respective owners. This ensures the evaluator cannot see the model weights, and Google DeepMind cannot see the evaluator’s test prompts.
The announcement highlights the risks of benchmark contamination, where models may have prior exposure to test questions, artificially inflating performance scores. Traditional safeguards like zero-logging protocols and contractual confidentiality have been used to mitigate this, but the new approach adds technical and cryptographic protections as a major step forward for secure model evaluation.
Google DeepMind emphasized that as AI models become more capable, ensuring models have not seen test questions in advance is critical for high-stakes evaluations, such as those used for cybersecurity or government assessments. The company stated that independent organizations need to trust that AI benchmarks accurately reflect a model’s true capabilities and safety, and double-blind evaluations aim to address this gap by eliminating the tradeoff between prompt confidentiality and model IP protection.
- Aug 26, 2026 · arXiv cs.AI
New NL2SQL benchmark shows enterprise database complexity degrades model performance
Trust79 - Aug 26, 2026 · arXiv cs.AI
New benchmark shows how reader-facing memory formats affect LLM evaluation scores
Trust79 - Aug 25, 2026 · arXiv cs.CL
Wazobia Eval introduces a 550-example Nigerian Pidgin benchmark for emotion, sarcasm, and cultural reasoning
Trust79