Skip to content
AImpact
IT EN
Models Intermediate Also known as: Mixture of Experts · Miscela di esperti

MoE

/em-oh-ee/

An architecture where the model is split into many specialized sub-models ('experts') and only a small share of them is activated for each token.

ShareLinkedInX

In practice

It enables models with hundreds of billions of parameters but inference cost closer to a much smaller one. Mixtral, DeepSeek, and GPT-4 use it. For API users nothing changes, but it explains surprising quality-to-price ratios.

Related terms

Seen in the wild

24 entries mentioning it
  1. Z.ai opens GLM-5.3-Flash: mystery model Ox Alpha revealed, MIT-licensed 320B MoE
    High
  2. Alibaba publishes Qwen3.8-Max weights: the first downloadable Max-class model
    High
  3. Moonshot AI ships Kimi K3: 2.8 trillion parameters, the largest open-weight model ever released
    High
  4. Thinking Machines Lab releases Inkling: a 975B-parameter multimodal MoE with Apache 2.0 open weights
    High
  5. Meta releases Llama 4.1: Scout, Maverick, and Behemoth MoE models under Apache 2.0
    Landmark
  6. DeepSeek V4 Preview: 1.6T parameters, 1M context, open weight in two sizes
    Landmark
  7. Alibaba releases Qwen 3.5: sparse multimodal MoE, 262K context, Apache 2.0
    High
  8. DeepSeek R2: the Chinese lab relaunches its open-weight reasoning model
    High
  9. Llama 4 Scout: 109B multimodal MoE with 10M context and vision SOTA
    High
  10. Qwen 3: Alibaba ships an open-weight family from 0.6B to 235B with native thinking
    High
  11. Llama 4: Meta moves to MoE and native multimodal, but the community is unimpressed
    High
  12. DeepSeek-V3-0324: the quiet update that puts vendor lock-in on notice
    Medium
  13. DeepSeek-V3: GPT-4o Quality at $0.55/M Tokens via MLA and FP8 Pipeline
    High
  14. DeepSeek-V3: China releases a shockingly cheap open frontier model
    Landmark
  15. DeepSeek-Coder-V2: GPT-4 Turbo coding quality with open weights
    High
  16. DeepSeek-V2: Multi-head Latent Attention and the first highly efficient Chinese open MoE
    High
  17. Mixtral 8x22B: Mistral's Apache 2.0 MoE with 39B active parameters
    High
  18. Snowflake Arctic: 480B total / 17B active MoE, enterprise SQL SOTA
    Medium
  19. DBRX: Databricks's 132B-total / 36B-active open MoE
    Medium
  20. Gemini 1.5 Pro: 1 million tokens in context
    High
← All terms