Meta Llama 3.2: First Multimodal Open Llama, 1B Runs on iPhone
In one sentence Meta releases Llama 3.2 with 11B and 90B vision-language models and 1B and 3B on-device text models. The 1B model runs on iPhone. Apache 2.0 license.
Meta released Llama 3.2, and this is a historic release: for the first time, the Llama family can see, not just read. The 11B and 90B models are vision-language, meaning they understand both text and images. Show them a photo and they can describe it, analyze a chart, or read text within an image. But perhaps the even bigger news is the small models: 1B and 3B parameters designed to run directly on-device — your phone or laptop — without needing the internet. The 1B model runs on an iPhone. This is a turning point: until now, having an open multimodal LLM required powerful servers; now a phone is enough. Everything is under Apache 2.0 license, meaning you can use it freely even in commercial products. With Ollama, setup is a single command. Meta proved that open source can match closed models on multimodal capabilities.
Companies
Meta
Tools
Llama 3.2, Ollama
Tags
Sources