Skip to content
Safety · Aug 18, 2026

Benchmark finds frontier LLMs leak sensitive memory data in up to 69% of tests

New CIMemories benchmark shows persistent memory in leading models produces high rates of inappropriate information disclosure, with violations rising with repeated use.

Trust79
HypeLow hype

2 sources · cross-referenced

ShareXLinkedInEmail
TL;DR
  • A new benchmark called CIMemories evaluates whether LLMs control information flow from persistent memory based on task context.
  • Frontier models exhibited up to 69% attribute-level violations, leaking sensitive data in inappropriate contexts.
  • Violations increased with repeated use: GPT-5’s violations rose from 0.1% after one task to 9.6% after 40 tasks.
  • Privacy-conscious prompting did not resolve the issue; models overgeneralized by sharing too much or too little information.
  • A separate method using explicit reasoning and reinforcement learning reduced inappropriate disclosures while maintaining task performance.

A new benchmark called CIMemories evaluates whether large language models (LLMs) appropriately control the flow of information from their persistent memory based on task context. The benchmark uses synthetic user profiles containing over 100 attributes per user, paired with diverse task contexts where each attribute may be essential for some tasks but inappropriate for others.

In evaluations of frontier models, researchers observed up to 69% attribute-level violations, where models inappropriately disclosed sensitive information from memory. Lower violation rates were often achieved at the expense of task utility, indicating a trade-off between privacy and performance.

Violations accumulated with repeated use: for GPT-5, violations increased from 0.1% after a single task to 9.6% after 40 tasks. When the same prompt was executed five times, violations reached 25.1%, revealing arbitrary and unstable behavior where models leaked different attributes for identical prompts.

Privacy-conscious prompting did not resolve the issue. Models tended to overgeneralize, either sharing too much information or too little, rather than making nuanced, context-dependent decisions about disclosure.

The researchers conclude that these findings reveal fundamental limitations in current models’ ability to reason contextually about information sharing. They argue that addressing these issues will require capabilities beyond better prompting or scaling, such as contextually aware reasoning mechanisms.

A separate paper proposes a method to improve contextual integrity by combining explicit reasoning with reinforcement learning. Using a synthetic dataset of 700 examples with diverse contexts and disclosure norms, the approach substantially reduced inappropriate information disclosure while maintaining task performance across multiple model sizes and families. Improvements transferred to established contextual integrity benchmarks like PrivacyLens, which uses human annotations to evaluate privacy leakage in AI assistants.

Sources
  1. 01Schneier on SecurityLLMs and Contextual Integrity
  2. 02arXivCIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
Also on Safety

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.