
Robust evaluation of vision models for real-world video tasks.
Vision Arena measures frontier vision models against benchmarks and evaluations that match the complexity of real-world video tasks.
2026/09/10Introducing Vision Arena
The Visual Intelligence Index
The Visual Intelligence Index combines different benchmarks to evaluate vision model capabilities in understanding video.
Visual Intelligence Index
Visual Intelligence Index v0 is a weighted average of 7 benchmarks. Updated September 11, 2026
13 of 13 models
Display
Capabilities Index Scores
Each index is a cross-benchmark weighted score of benchmark questions classified by capability.
13 of 13 models
Display
| Model | Video-MME v2 | LVBench | Perception Test | NExT-QA | Q-Bench Video | EgoSchema | UCF101-AD | Average |
|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra | - | - | 92.5% | 90.7% | - | 83.4% | 61.4% | 82.0% |
| Gemini 3.7 Flash | - | - | 85.2% | 88.7% | - | 81.0% | 59.6% | 78.6% |
| Gemini 3.5 Flash | - | - | 81.7% | 86.2% | - | 79.6% | 65.7% | 78.3% |
| Qwen3.8-Max | - | - | 87.0% | 88.7% | - | 84.8% | 51.0% | 77.9% |
| Gemini 3.8 Flash | - | - | 86.1% | - | - | 80.4% | 63.0% | 76.5% |
| Claude Fable 5.1 | - | - | 77.3% | 86.8% | - | 80.2% | 60.6% | 76.2% |
| Gemini 3.1 Pro | - | - | 80.2% | - | - | 78.2% | 68.5% | 75.6% |
| Gemini 3.6 Flash | - | - | 82.6% | - | - | 78.8% | 57.0% | 72.8% |
| Kimi K3 | - | - | 82.4% | 86.3% | - | 78.6% | 39.4% | 71.6% |
| Qwen3.8-27B | 50.5% | 64.6% | 72.3% | 82.2% | 66.7% | 79.1% | 45.5% | 65.8% |
| Qwen3.5-35B-A3B | 42.5% | 64.4% | 66.7% | 83.4% | 66.4% | 77.2% | 41.3% | 63.1% |
| Nemotron 3 Nano Omni 30B-A3B | 35.0% | 54.2% | 70.3% | 81.2% | 64.9% | 65.8% | 31.0% | 57.5% |
| Molmo 2 8B | 27.5% | 48.6% | 67.4% | 89.2% | 62.9% | 58.2% | 19.5% | 53.3% |
Benchmark scores report accuracy.
Vision Benchmarks
Video-MME v2
Full-spectrum video
A benchmark that scores video understanding through grouped questions, factoring in answer consistency and reasoning coherence.
LVBench
Long-context reasoning
A benchmark tailored for comprehensive long video understanding on hour-plus videos across sports, documentaries, events, TV shows, and more.
Perception Test
Spatial perception
A benchmark by DeepMind for low- and mid-level visual perception, covering memory, physics, and semantics across video, audio, and text.



