In practice
Claude and GPT-4 read images, Gemini handles video, some models talk in voice. For products this means analyzing receipt photos, screenshots, charts without a separate OCR. Watch out: visual input costs more tokens.
Related terms
Seen in the wild
30 entries mentioning it- LandmarkMeta releases Llama 4.1: Scout, Maverick, and Behemoth MoE models under Apache 2.0
- HighGoogle I/O 2026: Gemini Ultra 3, Project Astra goes live on Pixel, 2M context with real-time grounding, Veo 3.2, Imagen 4
- HighMistral Small 4: three models (reasoning + vision + coding) fused into one open weight
- MediumNano Banana 2: Google rebuilds its viral image model around consistency and text
- HighDeepSeek releases Janus Pro: one model to understand and generate images
- HighGemini 3 Pro and Flash: Google relaunches the frontier challenge
- LandmarkAlibaba releases Qwen2.5-VL 72B: best open-source multimodal model beats GPT-4o on key benchmarks
- HighOllama 1.0: first stable release with multimodal, tool calling, and Windows GA
- MediumOllama native vision model support: local VLMs with a one-liner
- HighKimi VL Thinking (Moonshot AI): first open visual model with RL-trained chain-of-thought reasoning
- HighLlama 4: Meta moves to MoE and native multimodal, but the community is unimpressed
- HighGemini 2.0 Flash Thinking: multimodal reasoning with visual chain-of-thought
- HighGemini 2.0 Flash GA: Google ships its fast multimodal model to production
- MediumSmolVLM2 (HuggingFace): 2.2B VLM for video and image understanding on consumer hardware
- LandmarkGemini 2.0 Flash: natively multimodal with audio and image output
- LandmarkGemini 2.0 Flash: Google opens the 'agentic era' and shows Astra/Mariner/Jules
- MediumPixtral: Mistral brings vision to European open models
- HighLlama 3.2: Meta brings vision and edge to open models
- MediumAgno (formerly Phidata): lightweight, multimodal agent framework 10x faster
- LandmarkGoogle Gemini 1.0: natively multimodal in three sizes