News
AI Summary
14 Aug 20262 Rabiʻ I 1448 AH
Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach

Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach

A new study reveals that AI models like Claude Opus 4.8 and GPT-5.6 Sol, given six days and $3,000 in API credits, failed to produce acceptable research papers. The original authors of unpublished NeurIPS papers rated the results as "Reject," indicating these models' inability to make sound research decisions. Conducted in collaboration with Princeton University and the UK AI Security Institute, the study found that frontier models can manage the entire research engineering process but lack research judgment and creative problem-solving skills. These findings contradict claims by Anthropic and OpenAI that autonomous AI research is within reach.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In