Ling-3.0-flash-VL offers downloadable image and video understanding
inclusionAI’s multimodal model accepts text, images, and video and produces text. Its official repository provides public weights.
Language & visionDownloadable weights
By Privaj Scout · Brief published Sep 16, 2026 · Sources checked Sep 16, 2026
SCOUT’S TAKE
Useful for investigating visual questions and document or video understanding. Its small active-parameter count does not describe its full memory footprint.
Total parameters
124B · maker-reported
Active per token
5.5B
Context
Up to 256K with documented configuration
License
MIT
What changed.
01
The maker reports 124B total parameters, with 5.5B active per token.
02
The documented extended-context configuration supports up to 256K tokens.
03
The official repository identifies the license as MIT.
The maker’s 256K recipe uses four 141GB-class GPUs. That is a documented configuration, not a minimum VRAM estimate or a home-computer speed test.
What the evidence shows
The date is repository creation, not a verified launch day. The Hub’s tensor inventory displays about 125B; the 124B architecture figure above is explicitly maker-reported.