News
AI Summary
18 Sept 20267 Rabiʻ II 1448 AH
Zhipu Opens GLM-5.3-FlashX Near 200 Tokens/s on ~100k Domestic Accelerators

Zhipu Opens GLM-5.3-FlashX Near 200 Tokens/s on ~100k Domestic Accelerators

Zhipu AI has announced the launch of GLM-5.3-FlashX on its API, claiming this version achieves inference speeds of up to 200 tokens per second. This speed relies on the inference capacity of around 100,000 domestic AI accelerators, along with further investments in infrastructure and service optimization. The base model, GLM-5.3-Flash, was open-sourced on August 26, featuring a 320-billion-parameter mixture-of-experts model with about 18 billion active parameters. The production serving stack for Flash was built from scratch on the domestic accelerator cluster, focusing on memory and bandwidth constraints.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In