News
AI Summary
20 Sept 20269 Rabiʻ II 1448 AH
DeepSeek-V4.1-Flash Ships Causal Encoder–Decoder MoE With 1M Context and Extreme KV Compression

DeepSeek-V4.1-Flash Ships Causal Encoder–Decoder MoE With 1M Context and Extreme KV Compression

DeepSeek has announced the launch of DeepSeek-V4.1-Flash, a multimodal model featuring 552 billion parameters. The model utilizes a Causal Encoder-Decoder design with a one-million-token context window, positioning it as the smallest member of a new architecture family. The model employs a 40-layer Transformer, split into a 20-layer causal encoder and a 20-layer decoder, allowing for efficient processing of heavy workloads. It was trained on approximately 45 trillion multimodal tokens, with performance enhancements through techniques like Compressed Sparse Attention 2.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In