Rank #8 / 13
Gemini 3.6 Flash
Gemini 3.6 Flash is Google's Gemini model (proprietary). It is ranked 8 of 13 models with enough coverage on the Visual Intelligence Index.
- Proprietary
- Gemini
- Video input
- Audio
- Structured output
- Tool calling
- Streaming
- Batch API
Comparison Summary
Gemini 3.6 Flash ranks 8 of 13 on the perception index. Immediately ahead is Claude Fable 5.1. Immediately behind is Gemini 3.1 Pro. List price is $0.35 per million input tokens and $2.80 per million output. 5 of 13 listed models cost less to prompt. Latency is 5.5s with a 1M context window and — tokens/s. Prices and most scores in this snapshot are placeholders, not sourced figures. Strongest capability in this set is Spatial Reasoning (83%). Weakest is Action / Event Understanding (69%).
- Visual Intelligence
- 79.6%
- Availability
- Proprietary
- Input / 1M tokens
- $0.35
- Output / 1M tokens
- $2.80
- Latency
- 5.5s
- Context Window
- 1M
- Speed (tokens/s)
- —
- Parameters
- —
- Time to first token
- —
All benchmark scores
| Benchmark | Metric | Score | 95% CI | Setting | Source |
|---|---|---|---|---|---|
| Video-MME v2 | Accuracy, no subtitles | — | — | — | — |
| LVBench | Test accuracy | — | — | — | — |
| Perception Test | Overall accuracy | 82.6% | 79.1–86.1 | 1 fps | Reported |
| NExT-QA | Accuracy (hard split) | — | — | — | — |
| Q-Bench Video | Overall accuracy | — | — | — | — |
| EgoSchema | Accuracy (fullset) | 78.8% | 75.2–82.4 | 1 fps | Reported |
| UCF101-AD | Accuracy | 57.0% | 55.5–58.5 | 1 fps | Reported |
Benchmark Scores
Compare reported model scores across each available benchmark or capability index.
13 of 13 models
Display
Capabilities Index Scores
Compare reported model scores across each available benchmark or capability index.
13 of 13 models
Display
Usability
4.0/ 5
What it's like to build against this model, scored out of five from the developer-facing capabilities in the dataset.
- Reachable1.0
There is an endpoint you can call without hosting anything.
- First-party API — supported
- Third-party API — supported
- Portable0.0
You can run it yourself, and are not tied to one vendor.
- Open weights — not supported
- Self-hostable — not supported
- Multimodal input1.0
It takes the footage directly, rather than frames you extracted.
- Video ingestion — supported
- Audio — supported
- Programmable1.0
Output you can parse, and tools it can call on its own.
- Structured output — supported
- Tool calling — supported
- Operable1.0
Usable interactively and in bulk, not only one call at a time.
- Streaming — supported
- Batch API — supported
Other Google models
Other models from the Google family.
| Date | Model | Visual Intelligence | Latency |
|---|---|---|---|
| Dec 4, 2025 | Gemini 3.1 Pro | 78.7% | 7.0s |
| Feb 18, 2026 | Gemini 3.5 Flash | 82.8% | 6.0s |
| Apr 30, 2026 | Gemini 3.6 Flash | 79.6% | 5.5s |
| Jul 9, 2026 | Gemini 3.7 Flash | 84.9% | 5.5s |
| Sep 2, 2026 | Gemini 3.8 Flash | 82.7% | 5.9s |