News
AI Summary
26 Jun 202611 Muharram 1448 AH
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run

An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run

Epoch AI has launched the new MirrorCode benchmark, which tests AI models' ability to recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56% solve rate, successfully rebuilding a 16,000-line toolkit in just 14 hours. The MirrorCode benchmark is crucial for assessing AI models' efficiency in complex programming tasks. Despite the impressive performance of Claude Opus 4.7, all tested models failed on the most complex tasks, highlighting ongoing challenges in the field.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In