August 19, 2026 · MIT Technology Review / Princeton et al.
AI Agents Ran Full Research Projects With a $3K Budget in 6 Days, but Both Papers Were Rejected
My take: This study says something many in the tech sector would rather not hear: AI cannot yet do open-ended science autonomously. Princeton, Stanford, UC Berkeley, and three other institutions gave Claude Opus 4.8 and GPT-5.6 Sol six days, $3,000 in API credits, and GPU access to tackle the core research questions of two NeurIPS 2026 papers. The agents completed all the engineering on their own, but the original authors rejected both papers (scoring them 2/6 and 1/6).
The distinction that emerges here is important: AI excels at research engineering tasks (running experiments, debugging code, processing data), but fails at the kind of open-ended, creative thinking that real science requires. The five failure modes identified include poor judgment about what is publishable, inability to improvise when a research design is not working, and a tendency to keep pushing in the same direction even when it is not yielding results.
For anyone using AI in their work, the practical takeaway is this: use AI where the task has clear success criteria. In areas where defining success is itself part of the challenge, AI remains a support tool, not a substitute. The question is: can you identify which of your daily tasks fall into each category?
Want to use these tools? See the unbiased reviews or back to the news.