News
AI Summary
27 Aug 202615 Rabiʻ I 1448 AH
Alibaba Open-Sources Qwen3.8-Flash with 6B Active Parameters at One-Ninth Training Cost

Alibaba Open-Sources Qwen3.8-Flash with 6B Active Parameters at One-Ninth Training Cost

On August 26, Alibaba's Qwen team released Qwen3.8-Flash, an open multimodal mixture-of-experts model. This release serves as a technical preview for the upcoming Qwen4 architecture, retaining a 3.x label while introducing a new generation for community validation. The main model features 125 billion parameters, with an additional 51 billion-parameter N-gram embedding module, activating only about 6 billion parameters per token. This new design, based on a mixture-of-experts approach, significantly reduces compute costs compared to dense models.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In