Introducing Vision Arena
Why build an arena?
OnDeck AI, Justin Chan · September 10, 2026
Visual reasoning is not solving a math word problem with pixels attached. It is solving a spatial and temporal puzzle: what is where, what changed, in what order, and when.
Open benchmarks and indexes
Real-world tasks, private benchmarks, more indexes
Arena: Elo-based voting
Video agents, world models, VLAs
Everything on this site is stage one. The three that follow are what we are building toward, not what we are claiming — nothing below describes a closed benchmark, a vote, or an agent evaluation, because none of them exist yet.
Vision Arena exists because the published numbers stopped predicting the work. A model can sit near the top of every video benchmark and still fail the first real task it is handed: find the moment the load shifted, count how many crates left the bay, say whether the person who came back is the person who left.
Math-in-an-image is a useful probe of diagram parsing. It is not a substitute for video intelligence. The hard problem is keeping a world coherent while it moves, then retrieving the moment that answers the question.
So this site does two things. It collects published results under a protocol strict enough that a comparison means something, and it says where every figure came from. Where a number cannot be sourced, the cell stays empty — it is never filled with a zero, and never averaged away.
The most common use case of Vision Language Models (VLMs) is for answering queries or extracting information from images. The other use case is for video analysis, but analyzing video is an extremely complex and compute-intensive task.
Over the years, frontier labs have released benchmarks to evaluate the ability of these vision models to understand video. However, despite improvements in benchmark scores, a clear disparity between practical results and evaluation performance has grown, indicating that the available benchmarks do not effectively replicate real video analysis workflows.
To come anywhere close to human-level analysis, models must possess the baseline set of capabilities outlined here.
Video understanding benchmarks typically assess a handful of these capabilities. Real vision tasks require all of them at once.
Vision Arena assesses models across a range of benchmarks to find which of them are capable of true video understanding rather than of one column of it.
"Visual question answering" is a format, not a skill. The same VQA wrapper can ask for a count, an identity, a spatial relation, a causal order, or a timestamp. Collapsing those into one accuracy hides the actual puzzle. We separate tasks along a few axes.
Traditional tasks are not "easy" and should not be treated as solved warm-ups. They are the perceptual primitives the temporal and spatial puzzles sit on. A model that cannot count returns, or cannot tell two similar objects apart, will look clever on a language exam and fail in footage.
The Visual Intelligence Index is the headline number on this site: one score per model, across every public video benchmark we track. It is computed, not judged.
This is a crude index, not a latent skill estimate. It is a summary of a table, and the table is the real artifact.
That is what the capability indexes are for. Long video, moment retrieval and identification are scored separately — filter the leaderboard by one and the ranks recompute for it, so a strength in one column cannot paper over a hole in another.
A useful video benchmark isolates a failure mode. Long-context suites ask whether evidence separated by minutes can still be combined. Temporal suites ask whether order, frequency and magnitude survive fluent language. Retrieval suites ask whether the model can return a timestamp, not a paragraph. Broad perception suites exist to stop the table collapsing onto one dataset's quirks.
We include a suite when it is well constructed, commonly cited for video understanding, large enough for a difference between models to mean something, and adds a length or a capability the rest of the set does not already cover.
A benchmark stops being helpful when it can be gamed by the input setting — more frames, subtitles, an extra retrieval tool — while the leaderboard presents those runs as comparable. Every published number here keeps its setting next to it.
See the benchmark directory for what each suite measures, and how each one is grouped into a capability index.
Evaluation is a protocol, not a screenshot of someone else's table. The comparable question is: same split, same metric, same input budget, no hidden transcript, no extra retrieval tool unless the column says so. If one paper reports 256 frames and another reports 1 fps, those are different tests. We record the setting and refuse to average it away.
This snapshot is a working dataset, not a finished measurement, and the site is built to show that rather than hide it.