AI agents lack creativity and judgment for open-ended research, study finds
Multi-institution research shows current AI agents can handle engineering tasks but fail to produce original, publishable work, tempering expectations for recursive self-improvement.
1 source · cross-referenced
- AI agents can perform engineering tasks like literature review and experimentation but struggle with creativity and judgment needed for original research.
- In a six-day test, agents produced papers rejected by original authors of NeurIPS 2026 submissions, failing to make novel contributions.
- Researchers propose 'shadow evaluation' method to assess open-ended AI research capabilities.
- Findings suggest timelines for automating AI research may be overstated.
A multi-institution research team led by Peter Kirgis and Sayash Kapoor at Princeton University evaluated whether current AI agents can conduct open-ended AI research—tasks requiring judgment, taste, and creativity rather than narrow engineering solutions. The study found that while AI agents can perform engineering tasks such as reviewing literature, running experiments, and compiling results, they lack the capacity to produce original research at the level expected by top-tier conferences.
To test these capabilities, the researchers developed a method called 'shadow evaluation,' where AI agents were tasked with answering research questions from two unpublished papers submitted to NeurIPS 2026. The agents were given six days, $3,000 in Anthropic API credits, GPU compute, virtual computers, and access to the open web. Despite these resources, the agents' outputs were graded by the original authors of the papers and found to be insufficient for publication.
The agents demonstrated competence in executing experiments and managing workflows but struggled with core research skills. They ran experiments on tiny synthetic datasets, failed to explore diverse ideas, committed to unpromising approaches early, and could not effectively backtrack or rethink their strategies. They also misused resources such as tokens, compute, and time, and did not incorporate feedback from subagents or external tools meaningfully.
The study suggests that the training regimes of current AI models, particularly reinforcement learning, are better suited to tasks with clear, automatable success metrics rather than open-ended research. 'It’s harder to create environments to train these models when the task itself is open-ended,' said Kapoor. The team is now testing Anthropic’s Mythos model, released in April, under the same framework.
The findings temper expectations about recursive self-improvement—the idea that AI systems will soon automate their own advancement with minimal human oversight. Recent industry claims, such as Anthropic’s blog post 'When AI Builds Itself' and OpenAI’s GPT-5.6 Sol assisting in post-training smaller models, highlight the perceived progress toward this goal. However, the study’s results indicate that current AI systems lack the creativity and intuitive judgment necessary for such open-ended innovation.
The research acknowledges limitations, including the small sample size of two papers and potential bias in grading by the original authors, who knew the submissions were AI-generated. Still, the authors argue that open-ended evaluations provide a richer test of capability than traditional benchmarks, offering a more realistic assessment of AI’s research potential.
- Aug 23, 2026 · Apple — Machine Learning Research
Apple ML paper proves P-completeness of Boolean query evaluation over inverted indices
Trust84 - Aug 22, 2026 · Ahead of AI — Sebastian Raschka
Researcher publishes 48-minute video explaining Anthropic’s Claude text watermarking mechanism
Trust75 - Aug 22, 2026 · Apple — Machine Learning Research
Apple researchers propose iterative pseudo-labeling to improve Mandarin-English code-switching speech recognition
Trust79