Skip to content

All Models

Rank #3 / 13

Gemini 3.7 Flash

Gemini 3.7 Flash is Google's Gemini model (proprietary). It is ranked 3 of 13 models with enough coverage on the Visual Intelligence Index.

  • Google
  • Proprietary
  • Gemini
  • Video input
  • Audio
  • Structured output
  • Tool calling
  • Streaming
  • Batch API
Read the report
78.6%Visual Intelligence Index · Rank #2 of 13

Capability profile

Comparison Summary

Gemini 3.7 Flash ranks 3 of 13 on the perception index. Immediately ahead is Qwen3.8-Max. Immediately behind is Gemini 3.5 Flash. List price is $0.40 per million input tokens and $3.00 per million output. 6 of 13 listed models cost less to prompt. Latency is 5.5s with a 1M context window and tokens/s. Prices and most scores in this snapshot are placeholders, not sourced figures. Strongest capability in this set is Spatial Reasoning (87%). Weakest is Action / Event Understanding (77%).

Visual Intelligence
84.9%
Availability
Proprietary
Input / 1M tokens
$0.40
Output / 1M tokens
$3.00
Latency
5.5s
Context Window
1M
Speed (tokens/s)
Parameters
Time to first token

All benchmark scores

BenchmarkMetricScore95% CISettingSource
Video-MME v2Accuracy, no subtitles
LVBenchTest accuracy
Perception TestOverall accuracy85.2%82.0–88.51 fpsReported
NExT-QAAccuracy (hard split)88.7%87.8–89.51 fpsReported
Q-Bench VideoOverall accuracy
EgoSchemaAccuracy (fullset)81.0%77.6–84.41 fpsReported
UCF101-ADAccuracy59.6%58.1–61.11 fpsReported

Benchmark Scores

Compare reported model scores across each available benchmark or capability index.

13 of 13 models
Filter by model access

Display

Capabilities Index Scores

Compare reported model scores across each available benchmark or capability index.

13 of 13 models
Filter by model access

Display

Usability

4.0/ 5

What it's like to build against this model, scored out of five from the developer-facing capabilities in the dataset.

Reachable1.0

There is an endpoint you can call without hosting anything.

  • First-party API — supported
  • Third-party API — supported
Portable0.0

You can run it yourself, and are not tied to one vendor.

  • Open weights — not supported
  • Self-hostable — not supported
Multimodal input1.0

It takes the footage directly, rather than frames you extracted.

  • Video ingestion — supported
  • Audio — supported
Programmable1.0

Output you can parse, and tools it can call on its own.

  • Structured output — supported
  • Tool calling — supported
Operable1.0

Usable interactively and in bulk, not only one call at a time.

  • Streaming — supported
  • Batch API — supported

Other Google models

Other models from the Google family.

DateModelVisual IntelligenceLatency
Dec 4, 2025Gemini 3.1 Pro78.7%7.0s
Feb 18, 2026Gemini 3.5 Flash82.8%6.0s
Apr 30, 2026Gemini 3.6 Flash79.6%5.5s
Jul 9, 2026Gemini 3.7 Flash84.9%5.5s
Sep 2, 2026Gemini 3.8 Flash82.7%5.9s