Ant Group’s InclusionAI lab has released Ling-3.0-flash-VL, a multimodal AI model that scores 25 on the Artificial Analysis Intelligence Index while activating just 5.5 billion of its 124 billion total parameters per inference.
The model, part of Ant Group’s Ling (Bailing) series, natively processes image, text, and video input. It’s been open-sourced under the MIT license, with BF16 and FP8 weights available on Hugging Face and ModelScope.
What makes this model different
The headline number, 25 on the Intelligence Index, matches the score achieved by Ling-3.0-flash, the text-only version that preceded it. Earlier evaluations actually showed the VL (vision-language) variant scoring 4 points higher than the text version, suggesting performance may vary depending on the benchmark suite applied.
The model houses 124 billion parameters in total but only activates 5.5 billion for any given inference call. This mixture-of-experts style design means the model can deliver strong results without requiring the full parameter set to fire every time a query comes in.
The model also features a context window stretching between 256K and 262K tokens, particularly relevant for enterprise use cases where large documents, lengthy video inputs, or complex multi-step agent workflows need to be handled without truncation.
Visual feedback and agentic capabilities
One of the more notable technical features is what Ant Group calls a “visual feedback closed-loop mechanism.” In practice, this means the model can evaluate its own visual outputs and self-correct, a capability that matters enormously for production-grade applications where errors compound quickly.
The practical applications Ant Group is targeting include image-to-web code generation, GUI agent functionalities, and medical report analysis.
API access is available through Ling Studio and partner platforms.
The open-weight strategy
Releasing a model of this caliber under the MIT license is a deliberate competitive move. The MIT license is about as permissive as open-source licensing gets: companies can use, modify, and commercialize the model with essentially no restrictions.
The availability of FP8 weights alongside the standard BF16 format is a practical detail worth noting. FP8 quantization reduces memory requirements roughly in half compared to BF16, making the model accessible to organizations running on more modest GPU infrastructure.
Where this fits in the competitive landscape
A score of 25 on the Artificial Analysis Intelligence Index places Ling-3.0-flash-VL in the conversation with other efficiency-focused multimodal models. Hitting 25 with only 5.5 billion active parameters highlights the efficiency gains achievable through sparse activation architectures.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
15








English (US) ·