Frontier AI agents fail to conduct open-ended research in shadow evaluations
Princeton-led study finds agents lack judgment, creativity, and adaptability in six-day paper-writing tasks despite substantial compute and API budgets.
1 source · cross-referenced
- Two shadow evaluations asked frontier AI agents to draft research papers matching unpublished studies, with up to six days and thousands of dollars in API credits.
- Expert authors of the original papers rejected both agent-generated papers, citing lack of judgment, poor data use, and failure to backtrack or creatively respond to feedback.
- Agents spent less than half their API budgets, ignored explicit instructions, and abandoned ambitious targets within a day, according to detailed logs reviewed by the authors.
- The team calls the method 'shadow evaluations' and plans regular follow-ups, noting current limitations and potential biases in their approach.
A Princeton-led team conducted two "shadow evaluations" in which frontier AI agents were tasked with drafting research papers that matched the main research questions of two unpublished studies, with six days of wall-clock time and thousands of dollars in API credits provided to each agent.
The original authors of the papers reviewed the agents’ outputs and unambiguously rejected both agent-generated papers, citing the agents’ lack of judgment for conducting open-ended research.
The authors reported that agents quickly rejected promising directions based on low-quality or synthetic data, lacked awareness of available resources (spending less than 50% of their API budgets despite encouragement to use them), and did not creatively respond to feedback, instead adding caveats or doubling down on unpromising directions.
Agents retired their most ambitious research targets within the first day and did not fundamentally shift their approach afterward, while also ignoring explicit instructions about exploration time, frequency of AI self-reviews, and strict paper-length limits.
The team analyzed over a hundred hours of agent logs and designed the method with input from collaborators, including some UK AISI coauthors, calling the approach 'shadow evaluations' because the agents shadowed studies they had not been trained on and could not access online.
The authors acknowledge limitations: expert reviewers knew the papers were AI-generated, the sample size was small (two papers), and their own priors on recursive self-improvement could influence interpretation, though they sought collaborators with diverse views and detailed their potential biases in the paper.
They conclude that current frontier agents struggle with open-ended research, suggesting that progress on verifiable tasks may not translate to broad recursive self-improvement, and plan to run regular shadow evaluations to track whether targeted training or scaffold improvements can overcome these limitations.
- Aug 5, 2026 · The Verge — AI
White House AI testing framework excludes open models, lacks risk definitions
Trust75 - Aug 4, 2026 · The Verge — AI
Texas orders audits for new data centers before grid connection
Trust75 - Aug 3, 2026 · The Verge — AI
EU transparency rules under AI Act take effect, requiring disclosure of AI interactions and deepfakes
Trust79