Skip to content
AImpact
IT EN
High Multimodal AI · 2 min read

Thinking Machines Lab releases Inkling: a 975B-parameter multimodal MoE with Apache 2.0 open weights

In one sentence Thinking Machines Lab's first model is a 975-billion-parameter open-weights MoE (41B active), natively multimodal, trained on 45 trillion tokens with a 1-million-token context: Apache 2.0 license, no metered API, fine-tuning via the Tinker platform.

Needs review Official source
ShareLinkedInX
Reading level

After more than a year of anticipation and record fundraising, on July 15 Thinking Machines Lab — the company founded by former OpenAI CTO Mira Murati — showed its hand: its first model is called Inkling and, surprisingly, it is not a paid service but a model with freely downloadable weights, under the most permissive license there is (Apache 2.0).

Inkling is enormous on paper (975 billion parameters) but frugal at runtime: thanks to its mixture-of-experts architecture, only 41 billion activate per request. And it is multimodal from birth: trained from the start on text, images, audio, and video, it is not a text model with vision bolted on afterwards. For smaller teams there is Inkling-Small, a reduced, cheaper-to-run version.

The commercial bet runs against the current: no metered API. Thinking Machines does not sell tokens but the Tinker platform, through which companies can retrain and customize the model on their own data. In short: the model is free; you pay to make it your own model.

For the open source world it is a milestone date: a top-tier American lab chooses open weights as its launch strategy, on ground so far dominated by Chinese labs (DeepSeek, Qwen, Kimi) and Meta. For companies it means running a frontier model in-house, on their own servers, without sending data to anyone.

Companies

Thinking Machines Lab

Tools

Inkling, Inkling-Small, Tinker

Tags

Open WeightsMixture of ExpertsThinking Machines LabFine-tuning

Sources