Skip to content
Vision Benchmarks

Vision Benchmarks

Evaluation benchmarks for video understanding models. These benchmarks are used to evaluate the performance of video understanding models on a variety of capabilities, including long context, temporal reasoning, perception, retrieval, and more.

Video-MME v2

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark that scores video understanding through grouped questions, factoring in answer consistency and reasoning coherence.

LVBench

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark tailored for comprehensive long video understanding on hour-plus videos across sports, documentaries, events, TV shows, and more.

Perception Test

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark by DeepMind for low- and mid-level visual perception, covering memory, physics, and semantics across video, audio, and text.

NExT-QA

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A video question-answering benchmark for causal and temporal reasoning about object interactions in everyday activities.

Q-Bench Video

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality.

EgoSchema

Perception, Spatial Reasoning, Temporal Reasoning, Action / Event Understanding, Causal Reasoning, Information Retrieval

A benchmark for video understanding using temporal certificate sets to measure the actual reasoning length a task requires.

UCF101-AD

Action / Event Understanding

A benchmark built to evaluate MLLM "capacity for denial," a measurement of model agreeableness and the tendency to hallucinate actions occurring.

Why these benchmarks

These are the public video-understanding benchmarks worth aggregating.

Sound construction
Annotation quality high enough that a wrong answer is the model’s fault rather than the label’s.
Commonly cited
Reputable and widely reported for video understanding, so a score here can be compared against what providers publish.
Large enough
Enough items that the difference between two models is a result rather than sampling noise.
Range
Between them they span clip lengths from seconds to hours, and capabilities from perception to temporal reasoning to retrieval.

Real-world task benchmarks

Coming soon

Public benchmarks saturate, get trained on, and are written to isolate a capability rather than to resemble anyone’s work. A high score in these benchmarks does not signal real-world capability, but rather that the model can do the thing the benchmark measures.

Closing that gap needs tasks drawn from real work, held closed, and scored against what a domain expert would have written. Drafts below; none are scored yet.

Describe corrosion and damage on the subsea pipeline for infrastructure inspections

Watch ROV survey video of a subsea pipeline and fill in the inspection report: corrosion level on the anodes, freespan and vibration, any UXO, or hazards that may compromise the pipeline. Whether it warrants a scheduled repair.

Build highlight reels for newsrooms from archival & live footage

Given a story brief, search across archival & the incoming feed for every shot that could relate to the segment, return results a producer can build the highlight reel from.

Identify anomalous maritime behaviour consistent with smuggling or illegal fishing

Watch the commercial vessel traffic (not the pleasure craft) for suspicious vessels and anomalous behaviour, such as repetitive rendezvous or behaviour consistent with illegal fishing.

Triage security alerts from CCTV footage

Watch a night of camera feeds from around the facility, and rank the events worth a guard’s attention from the four hundred that aren’t, describing each the way an incident log would.

Flagging dangerous driving in dashcam footage

Review dashcam footage and flag the moments a claims adjuster would call dangerous or aggressive driving, such as tailgating, driver on their phone, or running red lights.

Monitor offshore wind turbines for bird strikes

Detect and count every bird crossing the frame, and then identify the species across buoy cameras & turbine mounted cameras.