Anthropic researchers use Claude Mythos to probe cryptographic algorithms, publish prompts and eval
Claude Mythos Preview spent 60 hours (~$100,000 in API costs) attempting to find weaknesses in HAWK and a reduced-strength AES variant; researchers shared the prompts and announced a new benchmark, CryptanalysisBench.
2 sources · cross-referenced
- Claude Mythos Preview was used for 60 hours (~$100,000 in estimated API costs) to probe HAWK and a weaker AES variant for cryptographic weaknesses.
- Anthropic published the prompts used to steer the model toward harder research problems, including spelling errors.
- The work included a new evaluation suite, CryptanalysisBench, developed with ETH Zurich, Tel Aviv University, and University of Haifa.
- Anthropic noted the results did not have practical impact on current systems.
Anthropic researchers report using Claude Mythos Preview to attempt to discover mathematical flaws in the HAWK cryptographic scheme and a reduced-strength variant of AES. The company states that neither result had practical impact on today’s computer systems.
The effort ran for a total of 60 hours at an estimated API cost of about $100,000. Human oversight focused on encouraging the model to persist and aim for publishable findings rather than low-hanging fruit.
Anthropic published the exact prompts used to steer the model, including spelling errors, illustrating how researchers coaxed the model into sustained cryptanalysis rather than giving up prematurely.
As part of the project, Anthropic and academic partners at ETH Zurich, Tel Aviv University, and the University of Haifa introduced a new evaluation suite called CryptanalysisBench to assess LLMs’ ability to perform cryptanalysis tasks.
The prompts highlight iterative prompting strategies, such as pushing the model to target harder variants (e.g., AES-128 R7) and to avoid settling for incremental or trivial results.
- Jul 28, 2026 · OpenAI — News
OpenAI highlights use of coding agents in scientific computing workflows
Trust72 - Jul 28, 2026 · Simon Willison — everything
Moonshot releases Kimi K3, a 2.8T-parameter model with new commercial-use restrictions
Trust79 - Jul 27, 2026 · Google DeepMind — Blog
Google DeepMind releases three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Trust79