Grok 4.6 icon
Grok 4.6Gold Verified Icon
#3 in LLM models
4.4/5
« A powerful LLM model designed for coding, research, and long-running agentic tasks, featuring 500,000 tokens of context, tool calls, and fine-tuned reasoning effort. Available directly in Cursor and Grok Build »
Tool of the Week Aug 10 – 16, 2026
Verified Icon
Verified Tool
Freemium 205545
Traffic & popularity
Rank #7,368·+30.4%·Semrush
Rank#7,368
Estimated traffic317,602
AI VisibilityCheck on Semrush
6-month trend+30.4%1.7M · 6 mo
Data · Aug 26Powered bySemrush

Grok 4.6: five extra intelligence points in five weeks, and the price stays put

Grok 4.6 is xAI's frontier language model built for long-running agents, coding and knowledge work. It landed a mere thirty-five days after Grok 4.5 and jumped from 56 to 61 on the Artificial Analysis Intelligence Index while keeping the same API rates. Developers reach it through Cursor, Grok Build, the xAI API, OpenRouter, Vercel and Cloudflare. Consumer Grok apps were absent from the launch notes, so treat this release as one for builders.

Pros
  • 500,000-token context window
  • API rates unchanged from Grok 4.5
  • 1753 Elo on GDPval-AA knowledge tasks
  • Finishes tasks in half the turns Opus 5 takes
Cons
  • Trails GPT-5.6 Sol on terminal benchmarks
  • Developer platforms only, with no consumer app rollout
  • Cached-input rate crept up from Grok 4.5

Grok 4.6 keeps its focus through hours of agent work

Grok 4.6 carries a complex task across many steps, checks its own output and refines it before moving on. xAI, now operating under the SpaceXAI banner, gave it a longer supplemental training run built on agentic reinforcement learning and curated data filtered by automatic checks.

In practice, you hand it an unfamiliar topic, it researches the subject, sketches an application, writes the main functions and then polishes the result. Its 500,000-token window swallows an entire code repository in one go.

  • Text and image input, text output with no stated limit
  • Four reasoning-effort levels, including the new xhigh setting
  • Knowledge cutoff of February 1, 2026
Matthew Berman's first hands-on look at Grok 4.6
Grok 4.6 takes on a design challenge inside Cursor

Wins the office work, still chasing the terminal

On the Artificial Analysis composite index, Grok 4.6 scores 61, level with GPT-5.6 Sol and just behind Claude Opus 5 and Claude Fable 5. It posts the best marks in xAI's launch table on knowledge-work evaluations, legal analysis included, yet gives ground on the pure software-engineering rows. Funny twist for a model marketed as a coder, its strongest suit turns out to be analyst-style work.

A word of caution is fair here (the figures come from xAI and Artificial Analysis, and no full independent replication had circulated at publication time).

EvaluationGrok 4.6Best measured rival
AA Intelligence Index61Claude Opus 5, 63
GDPval-AA v2, Elo1753Claude Opus 5, 1861
DeepSWE v1.165.9%GPT-5.6 Sol, 73%
Terminal-Bench v3.026%GPT-5.6 Sol, 34.6%

Where to run Grok 4.6 and what it costs

Grok 4.6 is billed at $2 per million input tokens and $6 per million output tokens on the xAI API, with cached input at $0.50. Past 200,000 prompt tokens the rates double, and a faster variant runs at twice the standard price.

In Cursor the model ships on every subscription plan, a generous move among AI coding assistants, and Grok Build made it the default. To get people testing, xAI doubled the included usage in both environments during launch week. These numbers reflect launch day, so check xAI's published pricing grid before committing a budget.

Frequently asked questions

Is Grok 4.6 free?

No, Grok 4.6 has no free tier. The xAI API bills it at $2 per million input tokens and $6 per million output tokens, and it comes bundled with Cursor subscriptions on every plan. During launch week, xAI doubled the included usage in Cursor and Grok Build.

Grok 4.6 vs Claude Opus 5, which should you pick?

Claude Opus 5 stays ahead on raw intelligence, 63 versus 61 on the Artificial Analysis index. Grok 4.6 answers with economics, roughly $0.84 per completed task against $2.03 for Opus 5 per CodingFleet, in about half the conversation turns. Many teams route routine jobs to Grok 4.6 and escalate only the hardest calls.

How reliable is Grok 4.6 on factual answers?

Artificial Analysis measured a 65.7% non-hallucination rate for Grok 4.6. Put plainly, about one wrong answer in three arrives stated with full confidence instead of being flagged as uncertain. Any customer-facing or financial deployment should keep a verification layer between the model and the reader.

Can you use Grok 4.6 on grok.com or in the Grok app?

The launch announcement lists developer surfaces only, meaning Cursor, Grok Build, the xAI API, OpenRouter, Vercel and Cloudflare. Consumer Grok apps go unmentioned. If writing code is not your thing, a Cursor subscription is the quickest way in, since the model is included on every plan.

Verdict: For agent pipelines that chew through thousands of coding, research or document tasks a month, Grok 4.6 is the natural first benchmark, and budget-conscious teams will get the most mileage out of it.

★ Featured AI Tools ★
AI Alternatives for Grok 4.6
Paid
Paid
Free
Paid
Paid
Freemium
Free
Paid