News
AI Summary
16 May 202629 Dhuʻl-Qiʻdah 1447 AH
Latest open artifacts (#21): Open model bonanza! Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 & others. On CAISI's V4 assessment.

Latest open artifacts (#21): Open model bonanza! Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 & others. On CAISI's V4 assessment.

Open Labs, including DeepSeek, released new models this month. The Center for AI Standards and Innovation (CAISI) assessed these open models, noting that the gap between them and American models is widening over time. The report employed nine different criteria, revealing that DeepSeek V4 performed poorly in some areas, impacting its overall evaluation. CAISI utilizes an Elo calculation method based on Item Response Theory, commonly used for model comparisons. However, the gap between open and closed models does not present the complete picture, as evaluations rely on simplistic benchmark settings.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In