Grok 4.5 · API · High effort
The alt-frontier code smith — near-Opus agentic coding at a third the token bill.
SpaceXAI Grok 4.5 for fast, cheap frontier coding and terminal work when you want Claude/GPT rivals without $25–30 output rates.
Overall mean of attribute public scores after peer min-max (40–95 band within scored launch peer set).
- Endpoint
- grok-4.5
- Product
- SpaceXAI / xAI API
- Price
- $2 / $6 per MTok
- Context
- 500k tokens
Grok 4.5 API at high reasoning, 500k context, $2/$6 per MTok. Alt-frontier coding and agentic Build — strong cost/tempo story vs Claude Opus and GPT Sol.
Six stats · peer public
Peer public (main number / bar) ranks this Build inside the current scored peer set (≈40–95). Raw is the absolute evidence-fusion aggregate before min-max. A low peer score can still be a strong absolute model — see methodology.
When to choose
- Cost-efficient agentic coding (named harnesses)
- Terminal / multi-step engineering loops
- Greenfield apps and high-tempo iteration
- EU-only deployments until regional GA
- Trust-critical judgment without verification (54% Omniscience hall)
- Maximum absolute coding ceiling at any cost (Sol / Fable still lead Verified)
You want near-frontier coding and agentic work with strong price and speed.
Task proficiencies
Signature strengths and honest flaws
Delivers near-frontier agentic coding at roughly a third of the flagship token bill and with high tempo — strong value when you want Claude/GPT-class work without $25–30 output rates.
- Cost-efficient agentic coding on named harnesses
- High-tempo iteration and terminal engineering loops
Fast, willing starter on greenfield apps and multi-step engineering loops — good tempo for iterating a new build from scratch without heavy ceremony.
- Greenfield apps and prototypes
- High-iteration multi-step engineering
Weak Omniscience story (~54% hallucination pattern) — asserts trust-critical judgments confidently even when wrong, so unverified factual answers are risky.
- Trust-critical judgment without verification
- Factual claims outside provided context
Workaround · Require a verification step or citations for trust-critical claims; lean on it for coding and tempo, not unverified knowledge work.
Availability gaps in some regions (notably EU) until regional GA, which can block deployments with data-residency or availability requirements.
- EU-only or region-locked deployments
- Compliance requiring specific data residency
Workaround · Confirm regional availability before committing; keep a fallback Build for region-locked or EU-only workloads.
Traceability
Every factual claim cites an evidence row. Tap a chip to see the benchmark, exact config, source, and caveats behind the number.
Engine pass meth_v0_1_1 (peer min-max vs scored launch peer set). Codex overall 74 (B). Attribute public scores map raw aggregates into a 40–95 band within the peer set (methodology min-max). INT uses hierarchical fusion (GPQA primary / HLE secondary / capped bleed).