News
AI Summary
16 May 202629 Dhuʻl-Qiʻdah 1447 AH
New benchmark shows Claude Mythos and GPT-5.5 can develop real browser exploits autonomously

New benchmark shows Claude Mythos and GPT-5.5 can develop real browser exploits autonomously

Researchers at Carnegie Mellon University have developed a new benchmark to assess AI agents' ability to exploit real vulnerabilities in Google's V8 engine. The benchmark reveals that Mythos significantly outperforms GPT-5.5, albeit at a cost twelve times higher. This benchmark is crucial for evaluating AI's capabilities in creating real browser exploits, providing precise data on performance and cost. It highlights the substantial gap between different models in this domain.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In