News
AI Summary
17 May 20261 Dhuʻl-Hijjah 1447 AH
New math benchmark reveals AI models confidently solve problems that have no solution

New math benchmark reveals AI models confidently solve problems that have no solution

A consortium of 64 mathematicians has launched the new SOOHAK benchmark, which includes 439 handwritten tasks, of which 99 are unsolvable. Google's Gemini 3 Pro leads in solving research problems with a 30% success rate. However, no model has surpassed 50% in recognizing broken tasks. The SOOHAK benchmark reflects the gap between impressive results and the broader research skills lacking in AI systems. Improving performance requires more computing power, yet this does not aid models in recognizing that some problems have no solutions. The benchmark underscores the challenges AI faces in research domains.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In