Skip to content
AImpact
IT EN
Models Beginner Also known as: Multimodale

Multimodal

A model able to handle multiple input and output types together: text, images, audio, video. Not just reading but also generating multiple formats.

ShareLinkedInX

In practice

Claude and GPT-4 read images, Gemini handles video, some models talk in voice. For products this means analyzing receipt photos, screenshots, charts without a separate OCR. Watch out: visual input costs more tokens.

Related terms

Seen in the wild

30 entries mentioning it
  1. Meta releases Llama 4.1: Scout, Maverick, and Behemoth MoE models under Apache 2.0
    Landmark
  2. Google I/O 2026: Gemini Ultra 3, Project Astra goes live on Pixel, 2M context with real-time grounding, Veo 3.2, Imagen 4
    High
  3. Mistral Small 4: three models (reasoning + vision + coding) fused into one open weight
    High
  4. Nano Banana 2: Google rebuilds its viral image model around consistency and text
    Medium
  5. DeepSeek releases Janus Pro: one model to understand and generate images
    High
  6. Gemini 3 Pro and Flash: Google relaunches the frontier challenge
    High
  7. Alibaba releases Qwen2.5-VL 72B: best open-source multimodal model beats GPT-4o on key benchmarks
    Landmark
  8. Ollama 1.0: first stable release with multimodal, tool calling, and Windows GA
    High
  9. Ollama native vision model support: local VLMs with a one-liner
    Medium
  10. Kimi VL Thinking (Moonshot AI): first open visual model with RL-trained chain-of-thought reasoning
    High
  11. Llama 4: Meta moves to MoE and native multimodal, but the community is unimpressed
    High
  12. Gemini 2.0 Flash Thinking: multimodal reasoning with visual chain-of-thought
    High
  13. Gemini 2.0 Flash GA: Google ships its fast multimodal model to production
    High
  14. SmolVLM2 (HuggingFace): 2.2B VLM for video and image understanding on consumer hardware
    Medium
  15. Gemini 2.0 Flash: natively multimodal with audio and image output
    Landmark
  16. Gemini 2.0 Flash: Google opens the 'agentic era' and shows Astra/Mariner/Jules
    Landmark
  17. Pixtral: Mistral brings vision to European open models
    Medium
  18. Llama 3.2: Meta brings vision and edge to open models
    High
  19. Agno (formerly Phidata): lightweight, multimodal agent framework 10x faster
    Medium
  20. Google Gemini 1.0: natively multimodal in three sizes
    Landmark
← All terms