DeepSeek releases Janus Pro: one model to understand and generate images
In one sentence Janus Pro is a unified 7B-parameter multimodal model that both understands images and generates them from text, outperforming DALL-E 3 and Stable Diffusion 3 on the GenEval benchmark. Fully open source and runs locally.
Until now, if you wanted an AI that could both analyze an image and create a new one from a text description, you needed two separate models. DeepSeek changed that with Janus Pro: a single 7-billion-parameter model that does both — and does them well.
On the GenEval benchmark, which measures the quality of image generation, Janus Pro beats both OpenAI's DALL-E 3 and Stability AI's Stable Diffusion 3. But it is not just an image generator: it is also a strong vision-language model, capable of answering questions about photos, describing scenes, reading text in images, and reasoning about visual content.
The whole thing fits in 7 billion parameters — a manageable size for anyone with modern hardware. And what makes it even more compelling? It is completely open source. Anyone can download it, run it locally on their own machine without sending data to external servers, modify it, and adapt it to their own needs.
For researchers and developers this is a turning point: a unified model reduces the complexity of AI pipelines, lowers costs, and paves the way for advanced multimodal applications that previously required much heavier infrastructure.
Companies
DeepSeek
Tools
Janus Pro
Tags
Sources