Skip to content
Safety · Jul 29, 2026

New benchmark shows frontier LLMs discovering novel cryptanalytic attacks

CryptanalysisBench evaluates five leading models across 191 tasks, with Anthropic’s Mythos Preview uncovering previously unknown vulnerabilities in Hawk and reduced-round AES.

Trust79
HypeLow hype

2 sources · cross-referenced

ShareXLinkedInEmail
TL;DR
  • A new benchmark, CryptanalysisBench, assesses LLMs’ ability to perform mathematical cryptanalysis across 191 tasks spanning six cryptographic primitive families.
  • Five frontier models (Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, GLM-5.2) broke 65%–86% of Tier 1 schemes, 6–12 Tier-2 schemes at full strength, and 24–61 scaled-down variants.
  • Anthropic’s Mythos Preview used the benchmark to find novel vulnerabilities, including a key-recovery attack exploiting a design flaw in SpoC AEAD and an error in KINDI’s CCA-security proof.
  • The benchmark includes three tiers: known practical breaks, primitives with no known practical break (full strength and scaled-down), and a challenge set of production primitives.
  • Researchers release CryptanalysisBench to track AI’s cryptanalytic capabilities and stress-test cryptographic schemes before deployment.

A new benchmark called CryptanalysisBench evaluates the ability of large language models (LLMs) to perform mathematical cryptanalysis, a domain at the intersection of mathematical reasoning and cybersecurity. The benchmark consists of 191 tasks across six families of cryptographic primitives, primarily drawn from four NIST standardization competitions. These tasks are organized into three tiers: (i) primitives with known practical breaks, (ii) primitives with no known practical break evaluated at full strength and as scaled-down variants, and (iii) a challenge set of production primitives at the frontier of cryptanalysis.

Five frontier models—Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, and the open-weights GLM-5.2—were tested on CryptanalysisBench. The models broke 65%–86% of Tier 1 schemes, 6–12 Tier-2 schemes at full strength, and 24–61 scaled-down variants. Beyond reproducing known results, the models also produced novel cryptanalysis, including a key-recovery attack that exploits a design flaw in the SpoC AEAD and an error in KINDI’s published CCA-security proof, both previously unknown to the authors.

Anthropic used CryptanalysisBench to test its Mythos Preview model, discovering new vulnerabilities in the Hawk authenticated encryption scheme and reduced-round variants of AES. The benchmark is intended to serve as a tool for tracking whether AI cryptanalysis becomes a serious factor in the field and as a scaffold for stress-testing cryptographic schemes before deployment.

The authors emphasize that the attacks surfaced by the benchmark represent an early snapshot of a rapidly evolving frontier. They suggest that AI-driven cryptanalysis may soon match or exceed the published state of the art, highlighting the need for ongoing evaluation and adaptation in cryptographic research and practice.

Sources
  1. 01Schneier on SecurityMeasuring LLMs’ Ability to Perform Cryptanalysis
  2. 02arXivCryptanalysisBench: Can LLMs do Cryptanalysis?
Also on Safety

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.