News
AI Summary
28 Aug 202616 Rabiʻ I 1448 AH
AI benchmarks have a trust problem and Google wants to fix it

AI benchmarks have a trust problem and Google wants to fix it

Google DeepMind is conducting a double-blind evaluation of an advanced AI model for the first time. This pilot project, in collaboration with the Singapore AI Safety Institute, employs cryptographic protection through Confidential Space to prevent Google from accessing test questions and evaluators from seeing the model weights. The project utilizes the Gemini Flash Lite model and could establish a new standard for tamper-proof AI benchmarks. This initiative comes at a time when AI benchmarks face trust issues, highlighting the urgent need for reliable evaluations in the field.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In