News
AI Summary
2 Sept 202621 Rabiʻ I 1448 AH
OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

OpenAI has announced that its upcoming Astra model will be the first to carry a "critical" rating for cyber capabilities. The company plans to monitor the model's chain of thought, but this monitoring is deemed unreliable in reflecting the model's actual decisions. Reports indicate that Astra's new architecture pushes more of its thinking into unreadable areas, complicating the understanding of how the model makes decisions. This trend may weaken the safety net even as the model's capabilities increase.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In