Gemini 3.1 Pro · API · High thinking
The reasoning multimancer — 1M context, strong science, calibrated knowledge.
Gemini 3.1 Pro for deep research, long-context synthesis, and multimodal work when judgment and modalities matter as much as pure SWE points.
Overall mean of attribute public scores after peer min-max (40–95 band within scored launch peer set).
- Endpoint
- gemini-3.1-pro-preview
- Product
- Google AI / Gemini API
- Price
- $2 / $12 per MTok
- Context
- 1049k tokens
Google Gemini 3.1 Pro Preview API at high thinking, 1M context, multimodal (text/image/audio/video), $2/$12 per MTok (≤200k). Reasoning multimancer for research, long-context, and visual work.
Six stats · peer public
Peer public (main number / bar) ranks this Build inside the current scored peer set (≈40–95). Raw is the absolute evidence-fusion aggregate before min-max. A low peer score can still be a strong absolute model — see methodology.
When to choose
- Deep research and scientific reasoning (GPQA-class)
- Long-context multimodal analysis (1M + vision/audio/video)
- Calibrated knowledge work (strong Omniscience story)
- Absolute max independent agentic coding (Sol / Opus still lead mini-swe)
- Strict budget token farms (output $12; thinking burns tokens)
- Teams that need only open weights
The quest is research, multimodal, or long-context — not pure SWE ceiling at any price.
Task proficiencies
Signature strengths and honest flaws
Reads and synthesizes across a very large (1M-token) context window, pulling a coherent answer from long documents, sprawling codebases, or mixed-media corpora that overflow smaller windows.
- Deep research and long-document synthesis
- 1M-token corpora and long-context analysis
Reasons natively over vision, audio, and video alongside text, so image- and media-grounded tasks stay in one Build instead of bolting on a separate vision pipeline.
- Vision / audio / video analysis with text reasoning
- Multimodal document and UI understanding
Strong Omniscience / calibration story — comparatively better at signalling uncertainty and knowing what it does not know, which lowers confident-wrong risk on knowledge work.
- Calibrated knowledge and scientific reasoning
- Research where a flagged unknown beats a confident guess
Mid-pack on independent agentic SWE boards despite frontier reasoning — Sol and Opus still lead pure coding, so it is not the pick for a maximum coding ceiling.
- Maximum independent agentic coding at any price
- Named-harness SWE ceiling comparisons
Workaround · Keep Gemini for research, long-context, and multimodal work; escalate the hardest pure-coding tasks to Sol or Opus in the Party.
Reasoning / thinking mode consumes output tokens on deep chains (output $12 per MTok), making cost harder to predict on long analytical runs.
- Long deep-reasoning chains in thinking mode
- Strict budget token farms
Workaround · Cap thinking budget where the API allows, and reserve deep-reasoning mode for questions that actually need it rather than every call.
Traceability
Every factual claim cites an evidence row. Tap a chip to see the benchmark, exact config, source, and caveats behind the number.
Engine pass meth_v0_1_1 (peer min-max vs scored launch peer set). Codex overall 75.1 (B). Attribute public scores map raw aggregates into a 40–95 band within the peer set (methodology min-max). INT uses hierarchical fusion (GPQA primary / HLE secondary / capped bleed).