Skip to content
AImpact
IT EN
Models Beginner Also known as: Multimodale

Multimodal

A model able to handle multiple input and output types together: text, images, audio, video. Not just reading but also generating multiple formats.

ShareLinkedInX

In practice

Claude and GPT-4 read images, Gemini handles video, some models talk in voice. For products this means analyzing receipt photos, screenshots, charts without a separate OCR. Watch out: visual input costs more tokens.

Related terms

Seen in the wild

32 entries mentioning it
  1. Z.ai opens GLM-5.3-Flash: mystery model Ox Alpha revealed, MIT-licensed 320B MoE
    High
  2. Meta releases Llama 4.1: Scout, Maverick, and Behemoth MoE models under Apache 2.0
    Landmark
  3. Google I/O 2026: Gemini Ultra 3, Project Astra goes live on Pixel, 2M context with real-time grounding, Veo 3.2, Imagen 4
    High
  4. Mistral Small 4: three models (reasoning + vision + coding) fused into one open weight
    High
  5. Nano Banana 2: Google rebuilds its viral image model around consistency and text
    Medium
  6. Alibaba releases Qwen 3.5: sparse multimodal MoE, 262K context, Apache 2.0
    High
  7. DeepSeek releases Janus Pro: one model to understand and generate images
    High
  8. Gemini 3 Pro and Flash: Google relaunches the frontier challenge
    High
  9. Alibaba releases Qwen2.5-VL 72B: best open-source multimodal model beats GPT-4o on key benchmarks
    Landmark
  10. Ollama 1.0: first stable release with multimodal, tool calling, and Windows GA
    High
  11. Ollama native vision model support: local VLMs with a one-liner
    Medium
  12. Kimi VL Thinking (Moonshot AI): first open visual model with RL-trained chain-of-thought reasoning
    High
  13. Llama 4: Meta moves to MoE and native multimodal, but the community is unimpressed
    High
  14. Gemini 2.0 Flash Thinking: multimodal reasoning with visual chain-of-thought
    High
  15. Gemini 2.0 Flash GA: Google ships its fast multimodal model to production
    High
  16. SmolVLM2 (HuggingFace): 2.2B VLM for video and image understanding on consumer hardware
    Medium
  17. Gemini 2.0 Flash: natively multimodal with audio and image output
    Landmark
  18. Gemini 2.0 Flash: Google opens the 'agentic era' and shows Astra/Mariner/Jules
    Landmark
  19. Pixtral: Mistral brings vision to European open models
    Medium
  20. Llama 3.2: Meta brings vision and edge to open models
    High
← All terms