DeepSeek tests V4-Flash-Vision-Exp: multimodal agents at Flash pricing
In one sentence On August 21 DeepSeek shipped the experimental V4-Flash-Vision-Exp model via API, a multimodal variant of V4-Flash that adds vision: agentic capabilities on images and screenshots claimed close to Anthropic Opus 4.8, while keeping text parity with V4-Flash.
DeepSeek, the Chinese lab known for powerful models at rock-bottom prices, made an experimental model called V4-Flash-Vision-Exp available through its API. The news: its budget model for code and agents can now see images, so it can work on screenshots, charts and user interfaces, not just text.
The interesting part is what it does with those images. It does not merely describe them: it uses them to act. An agent that sees the screen can figure out where to click, read a chart to answer a question, or fix an interface by looking at it, edging closer to what far more expensive Western flagship models do.
Bloomberg framed the release as a test model built to rival Anthropic Opus 4.8 on multimodal agent tasks. It is another beat in the script DeepSeek has repeated for two years: take a premium-tier capability and push it into the budget tier, forcing competitors to rethink pricing. Being an experimental version, though, it is too early to treat it as stable for production use.
Companies
DeepSeek
Tools
DeepSeek-V4-Flash-Vision-Exp
Tags
Sources