MQ
MODELQUEST
choose the right AI for the quest
← The Codex
MultimancerSagePeer public scoresmeth_v0_1_1as of 2026-07-11

Gemini 3.1 Pro · API · High thinking

The reasoning multimancer — 1M context, strong science, calibrated knowledge.

Gemini 3.1 Pro for deep research, long-context synthesis, and multimodal work when judgment and modalities matter as much as pure SWE points.

SharePost on X
Codex rating
75.1
Confidence B

Overall mean of attribute public scores after peer min-max (40–95 band within scored launch peer set).

Loadout
Endpoint
gemini-3.1-pro-preview
Product
Google AI / Gemini API
Price
$2 / $12 per MTok
Context
1049k tokens

Google Gemini 3.1 Pro Preview API at high thinking, 1M context, multimodal (text/image/audio/video), $2/$12 per MTok (≤200k). Reasoning multimancer for research, long-context, and visual work.

Core attributes

Six stats · peer public

Peer publicbars = rank in scored peer set (≈40–95) · raw = absolute fusion under the bar
STRExecution
46.8raw 76.4B39.853.8
DEXTempo
95.0raw 78.0B85100
CONReliability
42.2raw 76.4B35.249.2
INTReasoning
92.8raw 88.1A88.896.8
WISJudgment
95.1raw 81.3A91.199.1
CHACollaboration
78.6raw 85.6A74.682.6

Peer public (main number / bar) ranks this Build inside the current scored peer set (≈40–95). Raw is the absolute evidence-fusion aggregate before min-max. A low peer score can still be a strong absolute model — see methodology.

Practical verdict

When to choose

Best at
  • Deep research and scientific reasoning (GPQA-class)
  • Long-context multimodal analysis (1M + vision/audio/video)
  • Calibrated knowledge work (strong Omniscience story)
Avoid for
  • Absolute max independent agentic coding (Sol / Opus still lead mini-swe)
  • Strict budget token farms (output $12; thinking burns tokens)
  • Teams that need only open weights
Choose this when

The quest is research, multimodal, or long-context — not pure SWE ceiling at any price.

Skills

Task proficiencies

Creative Writing89.0 B
Professional Writing89.0 B
Quantitative Reasoning87.2 A
Following Complex Instructions84.9 B
Research Synthesis84.6 A
Frontend And Ui81.6 A
Visual Understanding81.6 A
Debugging80.1 B
Existing Codebase Work80.1 B
Refactoring80.1 B
Test Writing80.1 B
Fast Interactive Work78.0 B
Devops And Terminal68.5 B
Long Autonomous Runs68.5 B
Tool Calling68.5 B
Traits & Quirks

Signature strengths and honest flaws

Traits · 3
Long-Context Synthesistconf Binferred

Reads and synthesizes across a very large (1M-token) context window, pulling a coherent answer from long documents, sprawling codebases, or mixed-media corpora that overflow smaller windows.

Triggers when
  • Deep research and long-document synthesis
  • 1M-token corpora and long-context analysis
Multimodal Nativeconf Binferred

Reasons natively over vision, audio, and video alongside text, so image- and media-grounded tasks stay in one Build instead of bolting on a separate vision pipeline.

Triggers when
  • Vision / audio / video analysis with text reasoning
  • Multimodal document and UI understanding
Calibrated Knowledgeconf Cinferred

Strong Omniscience / calibration story — comparatively better at signalling uncertainty and knowing what it does not know, which lowers confident-wrong risk on knowledge work.

Triggers when
  • Calibrated knowledge and scientific reasoning
  • Research where a flagged unknown beats a confident guess
Flaws · 2
Coding Dipconf Binferredmoderatecommon

Mid-pack on independent agentic SWE boards despite frontier reasoning — Sol and Opus still lead pure coding, so it is not the pick for a maximum coding ceiling.

Shows up when
  • Maximum independent agentic coding at any price
  • Named-harness SWE ceiling comparisons

Workaround · Keep Gemini for research, long-context, and multimodal work; escalate the hardest pure-coding tasks to Sol or Opus in the Party.

Thinking Token Burnconf Cinferredmoderateoccasional

Reasoning / thinking mode consumes output tokens on deep chains (output $12 per MTok), making cost harder to predict on long analytical runs.

Shows up when
  • Long deep-reasoning chains in thinking mode
  • Strict budget token farms

Workaround · Cap thinking budget where the API allows, and reserve deep-reasoning mode for questions that actually need it rather than every call.

Evidence Terminal

Traceability

Every factual claim cites an evidence row. Tap a chip to see the benchmark, exact config, source, and caveats behind the number.

Engine pass meth_v0_1_1 (peer min-max vs scored launch peer set). Codex overall 75.1 (B). Attribute public scores map raw aggregates into a 40–95 band within the peer set (methodology min-max). INT uses hierarchical fusion (GPQA primary / HLE secondary / capped bleed).