Comparison of vision-language models on the Visual Intelligence Index, a weighted average across every benchmark we track, alongside list price per million tokens, latency, context window, and provider.
Rank
Model
Visual Intelligence
Input $/1M
Output $/1M
Latency
Context window
Provider
Open model page
01
GPT-6 Astra
88.7%
$10.00
$50.00
4.3s
1.1M
OpenAI
02
Qwen3.8-Max
85.6%
$2.00
$6.00
10.8s
1M
Alibaba
03
Gemini 3.7 Flash
84.9%
$0.40
$3.00
5.5s
1M
Google
04
Gemini 3.5 Flash
82.8%
$0.30
$2.50
6.0s
1M
Google
05
Gemini 3.8 Flash
82.7%
$0.75
$3.75
5.9s
1M
Google
06
Kimi K3
81.5%
$0.60
$2.50
8.2s
256K
Moonshot
07
Claude Fable 5.1
81.5%
$10.00
$50.00
5.1s
1M
Anthropic
08
Gemini 3.6 Flash
79.6%
$0.35
$2.80
5.5s
1M
Google
09
Gemini 3.1 Pro
78.7%
$1.25
$10.00
7.0s
1M
Google
10
Qwen3.8-27B
67.5%
$0.25
$1.00
15.7s
256K
Alibaba
11
Qwen3.5-35B-A3B
64.7%
$0.20
$0.90
9.4s
256K
Alibaba
12
Nemotron 3 Nano Omni 30B-A3B
59.8%
$0.15
$0.60
5.9s
128K
NVIDIA
13
Molmo 2 8B
57.5%
$0.10
$0.40
0.2s
128K
Ai2
Why these models
Leading models with published evidence of video-understanding performance in their last few releases. That last part does most of the filtering: plenty of capable models have never reported a video benchmark, and a leaderboard entry with nothing to corroborate it is a claim, not a measurement.