Skip to content
Vision Models

Vision Models

Comparison of vision-language models on the Visual Intelligence Index, a weighted average across every benchmark we track, alongside list price per million tokens, latency, context window, and provider.

RankModelVisual IntelligenceInput $/1MOutput $/1MLatencyContext windowProviderOpen model page
01GPT-6 Astra88.7%$10.00$50.004.3s1.1MOpenAI
02Qwen3.8-Max85.6%$2.00$6.0010.8s1MAlibaba
03Gemini 3.7 Flash84.9%$0.40$3.005.5s1MGoogle
04Gemini 3.5 Flash82.8%$0.30$2.506.0s1MGoogle
05Gemini 3.8 Flash82.7%$0.75$3.755.9s1MGoogle
06Kimi K381.5%$0.60$2.508.2s256KMoonshot
07Claude Fable 5.181.5%$10.00$50.005.1s1MAnthropic
08Gemini 3.6 Flash79.6%$0.35$2.805.5s1MGoogle
09Gemini 3.1 Pro78.7%$1.25$10.007.0s1MGoogle
10Qwen3.8-27B67.5%$0.25$1.0015.7s256KAlibaba
11Qwen3.5-35B-A3B64.7%$0.20$0.909.4s256KAlibaba
12Nemotron 3 Nano Omni 30B-A3B59.8%$0.15$0.605.9s128KNVIDIA
13Molmo 2 8B57.5%$0.10$0.400.2s128KAi2

Why these models

Leading models with published evidence of video-understanding performance in their last few releases. That last part does most of the filtering: plenty of capable models have never reported a video benchmark, and a leaderboard entry with nothing to corroborate it is a claim, not a measurement.