Mistral Releases Pixtral 12B: Multimodal Model That Runs on Consumer GPUs
In one sentence Pixtral 12B is Mistral's first vision-language model, handling multiple images and charts under Apache 2.0, runnable on a single consumer GPU.
Mistral has been known for making efficient, open-source text models. Pixtral 12B is their first step into multimodal territory — a model that understands both text and images. The key thing about Pixtral is that it can handle multiple images in a single prompt, so you can give it several charts, diagrams, or photos at once and ask questions that span all of them. It's particularly good at understanding charts, tables, and document layouts, not just photos of things. At 12 billion parameters, it runs on a single consumer GPU — an RTX 3090 or 4090 gets the job done. It's released under the Apache 2.0 license, meaning you can use it commercially without restrictions. Ollama support means you can be up and running locally in a few minutes. This democratizes document-understanding AI: tasks that used to require expensive cloud APIs can now run on hardware you already own.
Companies
Mistral AI
Tools
Pixtral 12B, Ollama
Tags
Sources