Alibaba releases Qwen2.5-VL 72B: best open-source multimodal model beats GPT-4o on key benchmarks
In one sentence Alibaba releases Qwen2.5-VL 72B under Apache 2.0, surpassing GPT-4o on multiple multimodal benchmarks with support for documents, charts, 20+ minute videos, multilingual OCR, and GUI agent actions.
Alibaba has released Qwen2.5-VL 72B, a powerful AI model that can see and understand images, videos, documents, and application screenshots. The big deal here is that it is fully open source under the Apache 2.0 license — meaning anyone can download, use, and build on top of it for free.
What can it actually do? A lot. It reads scanned documents in many languages, interprets charts and tables, analyzes videos up to 20 minutes long, and can even interact with graphical interfaces the way a human would — identifying buttons, filling in fields, and navigating screens.
The most striking result is the benchmark comparison: on several standard industry tests, Qwen2.5-VL 72B outperforms OpenAI's GPT-4o, which had long been considered the gold standard for multimodal AI.
For developers using tools like Ollama on their own machines, the model is available to run locally — though the 72 billion parameters require serious hardware, typically one or more GPUs with 40 to 80 GB of VRAM combined.
This release marks a turning point for open-source AI. For the first time, a free and open model can genuinely compete with — and in some areas surpass — the best paid offerings from major American tech companies. It shifts the calculus for any organization deciding whether to use cloud AI APIs or run their own models.
Companies
Alibaba
Tools
Qwen2.5-VL, Ollama
Tags
Sources