Official xAI Grok 4.7 announcement social image

xAI Grok 4.7: Same Price, Longer Coding Horizons

xAI (official pages also brand as SpaceXAI) released Grok 4.7, a same-price upgrade over Grok 4.6 aimed at longer coding and knowledge-work sessions. Day-one surfaces named in the launch post are Cursor, Grok Build, the Grok API, plus third-party coding harnesses, model routers, and cloud platforms. Consumer web, mobile, and Grok-on-X access are planned later, per the model card.

The commercial shape did not change: list pricing still opens at $2 / $6 per million input/output tokens under 200k prompt tokens. The product bet is persistence — a larger base, a longer reinforcement-learning run on multi-hour tasks, stronger self-checking, and native understanding of the Grok Bot harness — not a new price tier. Primary sources: Introducing Grok 4.7 · model docs · model card (PDF).

Image credit: xAI

Key points

  • What shipped: Grok 4.7 (grok-4.7) for coding, agentic work, and knowledge work; same list price and speed class as Grok 4.6.
  • Where: Cursor, Grok Build, Grok API, third-party harnesses/routers/clouds today; consumer surfaces later.
  • Pricing (<200k): $2 input / $0.50 cached / $6 output per 1M tokens. (≥200k): whole request bills at $4 / $1 / $12. Launch post mentions a fast variant at 2× price / 2× output speed; public docs currently list grok-4.7 only.
  • Context / modalities: 500,000-token context; text + image in, text out. Reasoning efforts: low / medium / high (default) / xhigh.
  • Vendor benches (self-reported): CursorBench 4.0 46.3% (xhigh) vs 4.6 40.4% (high); DeepSWE v1.1 71.0% vs 65.2% (high); Terminal-Bench 4.0 38.0% (xhigh, Grok Build harness) vs 20.3% (high). Fable 5.1 still leads several of the same charts at much higher list prices.
  • Caveats: Scores are vendor-reported and often cross-effort or harness-coupled. Absolute Harvey Legal scores remain low. Model card: not for unattended high-stakes medicine, law, finance, or safety-critical decisions. 9to6AI has not hands-on audited coding quality.

What shipped

xAI frames 4.7 as its most capable Grok for coding and knowledge work: longer trajectories on hard tasks, more careful self-verification, and a new safeguard stack. Training notes in the launch materials stress a larger base than 4.6 and a longer RL pass weighted toward problems that take many hours.

Public docs put the model at a 500,000-token context window, text + image input with text output, and reasoning efforts of low / medium / high (default) / xhigh. Rate limits on the docs page: 150 requests per second and 50M tokens per minute across us-east-1, us-west-2, and us-central-1.

Pricing from the live docs page matches 4.6. Cross the 200k prompt cliff and every token in that request bills at the higher tier — the cost detail most teams miss on long agent loops. Official announcement: x.ai/news/grok-4-7. Model page: docs.x.ai/docs/models/grok-4.7.

What changed vs Grok 4.6 and the frontier pack

Same bill, harder diet. Migrating from 4.6 is a model-id swap, not a budget renegotiation. The benches that move most are the ones that punish short attention: CursorBench 4.0, Terminal-Bench 4.0, and related long-horizon coding suites in the card.

Read the small print on comparisons. Several headline deltas compare 4.7 at xhigh to 4.6 at high. Terminal-Bench for Grok uses the Grok Build harness. CursorBench 4.0 is not comparable to older CursorBench versions. Treat peer tables as directional, then re-measure on your harness.

Price-performance vs Fable / Sol. xAI’s own comparison table puts Grok at $2/$6 against GPT-5.6 Sol at $4/$20 and Claude Fable 5.1 at $10/$50. On that same table, Fable still leads CursorBench 4.0 (51.8%) and Terminal-Bench 4.0 (57.9%). The pitch is not “beats every frontier chart” — it is “competitive enough, much cheaper, live in Cursor today.”

Knowledge work and safety claims. Launch materials show gains on AA Briefcase / GDPval-style office work and a rebuilt refusal/jailbreak stack. Capability does not equal permission to run unattended legal or clinical agents — absolute Harvey Legal scores remain in the teens, and the model card requires human oversight in high-stakes domains.

Community signal. Early HN discussion clusters around plain-English tone vs “Claudish,” Cursor day-one access, cache pricing ($0.50/M under 200k), and skepticism that vendor Terminal-Bench / token-efficiency stories will hold under independent harnesses. Useful color; not a substitute for your evals.

What it means for builders

If you are already on Grok 4.6. Flip to grok-4.7 when multi-hour coding, terminal agents, or marathon-style streaks matter. Re-measure token-per-task at high and xhigh — longer persistence can raise output tokens even when list price is flat.

If you live in Cursor or Grok Build. This is the practical day-one path. Consumer Grok apps are not the launch surface for 4.7 yet.

If you buy on price-performance. Grok remains the aggressive undercut vs Anthropic/OpenAI list rates for Sol/Fable-class work. Factor cache reads and the 200k cliff into agent budgets, not just the $2/$6 headline.

When to wait or dual-run. Regulated legal/clinical workflows still need human review. If your bottleneck is peak coding quality regardless of price, keep Fable/Astra (or your current winner) as primary and use Grok as a second reviewer — a pattern several HN commenters already describe.

What to watch

  • Independent Artificial Analysis / third-party harness numbers once they settle, especially Terminal-Bench and token-per-task vs 4.6.
  • Whether a separate fast API id appears in docs (mentioned in the launch post, not listed as its own model page yet).
  • Consumer rollout to web, mobile, and X.
  • Real Cursor Ultra / Grok subscription token burn at default vs xhigh effort.

Sources

Leave a Comment

Your email address will not be published. Required fields are marked *