News
AI Summary
10 Sept 202629 Rabiʻ I 1448 AH
Ant Group Opens Ling-3.0-flash-VL Multimodal Model With Vision Feedback Loop

Ant Group Opens Ling-3.0-flash-VL Multimodal Model With Vision Feedback Loop

Ant Group has launched the Ling-3.0-flash-VL model, the first multimodal model in its Ling series, through the inclusionAI organization on Hugging Face and ModelScope. Built on the Ling-3.0-flash architecture, it features 124 billion parameters, activating about 5.5 billion per token, and accepts inputs from images, text, and video. The model's design includes a vision feedback loop that treats seeing as part of an ongoing work cycle, making it suitable for GUI automation and medical report generation. In image-to-webpage tests, it demonstrated the ability to recover layout and component relationships, surpassing GPT-5.4's performance.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In