News
AI Summary
1 Sept 202620 Rabiʻ I 1448 AH
DeepSeek Open-Sources V4-Flash-Vision-Exp, Its First Native Vision Model

DeepSeek Open-Sources V4-Flash-Vision-Exp, Its First Native Vision Model

DeepSeek launched its first multimodal model, DeepSeek-V4-Flash-Vision-Exp, on Hugging Face on August 31 under an MIT license. This release marks the initial support for image input within the V4 series, packaging model files, a tokenizer, a prompt-encoding reference, and a minimal PyTorch inference implementation. The new model can describe images, read text in screenshots, and analyze charts, accepting JPEG, PNG, GIF, and WebP formats. According to DeepSeek, its performance on pure-text tasks remains on par with the V4-Flash release, while it gains substantial improvements on agent benchmarks requiring visual understanding.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In