Skip to content
Research · Aug 23, 2026

AI agents lack creativity and judgment for open-ended research, study finds

Multi-institution research shows current AI agents can handle engineering tasks but fail to produce original, publishable work, tempering expectations for recursive self-improvement.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • AI agents can perform engineering tasks like literature review and experimentation but struggle with creativity and judgment needed for original research.
  • In a six-day test, agents produced papers rejected by original authors of NeurIPS 2026 submissions, failing to make novel contributions.
  • Researchers propose 'shadow evaluation' method to assess open-ended AI research capabilities.
  • Findings suggest timelines for automating AI research may be overstated.

A multi-institution research team led by Peter Kirgis and Sayash Kapoor at Princeton University evaluated whether current AI agents can conduct open-ended AI research—tasks requiring judgment, taste, and creativity rather than narrow engineering solutions. The study found that while AI agents can perform engineering tasks such as reviewing literature, running experiments, and compiling results, they lack the capacity to produce original research at the level expected by top-tier conferences.

To test these capabilities, the researchers developed a method called 'shadow evaluation,' where AI agents were tasked with answering research questions from two unpublished papers submitted to NeurIPS 2026. The agents were given six days, $3,000 in Anthropic API credits, GPU compute, virtual computers, and access to the open web. Despite these resources, the agents' outputs were graded by the original authors of the papers and found to be insufficient for publication.

The agents demonstrated competence in executing experiments and managing workflows but struggled with core research skills. They ran experiments on tiny synthetic datasets, failed to explore diverse ideas, committed to unpromising approaches early, and could not effectively backtrack or rethink their strategies. They also misused resources such as tokens, compute, and time, and did not incorporate feedback from subagents or external tools meaningfully.

The study suggests that the training regimes of current AI models, particularly reinforcement learning, are better suited to tasks with clear, automatable success metrics rather than open-ended research. 'It’s harder to create environments to train these models when the task itself is open-ended,' said Kapoor. The team is now testing Anthropic’s Mythos model, released in April, under the same framework.

The findings temper expectations about recursive self-improvement—the idea that AI systems will soon automate their own advancement with minimal human oversight. Recent industry claims, such as Anthropic’s blog post 'When AI Builds Itself' and OpenAI’s GPT-5.6 Sol assisting in post-training smaller models, highlight the perceived progress toward this goal. However, the study’s results indicate that current AI systems lack the creativity and intuitive judgment necessary for such open-ended innovation.

The research acknowledges limitations, including the small sample size of two papers and potential bias in grading by the original authors, who knew the submissions were AI-generated. Still, the authors argue that open-ended evaluations provide a richer test of capability than traditional benchmarks, offering a more realistic assessment of AI’s research potential.

Sources
  1. 01MIT Technology Review — AIAI’s recursive self-improvement might not come so quickly after all
Also on Research

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.