Grok 4.6 Is Here — xAI's Strongest Model Yet, Now on Multos
xAI released Grok 4.6 today. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index at 61 points, jumps from 54.0% to 65.9% on DeepSWE, adds a new xhigh reasoning level, and keeps the same $2/$6 pricing as Grok 4.5. It's live on Multos now.
xAI shipped Grok 4.6 on August 12, 2026. The headline: it matches GPT-5.6 Sol Max on the Artificial Analysis Intelligence Index at 61 points — a 5-point jump over Grok 4.5's score of 56. DeepSWE climbed from 54.0% to 65.9%. Terminal-Bench went from 15.7% to 26.0%. A new xhigh reasoning level arrives on top of the existing low/medium/high settings. Price and context window are unchanged: $2/$6 per million tokens, 500K token context. It's been live on Multos since release.
What Changed from Grok 4.5
Same 1.5-trillion-parameter V9 architecture, same 500K context window, same $2/$6 pricing. What changed is everything inside the training pipeline. xAI ran a longer supplemental training pass with curated model-generated data for reasoning and advanced technical concepts, regenerated all SFT trajectories using Grok 4.5 itself as the teacher model, and trained on a wider range of agentic RL environments: knowledge work, general coding, kernel optimisation, web development, computer-aided design. The result is a model that holds its focus across longer sequences of tool calls and self-verifies more often before moving on — the company reports noticeably more self-testing behaviour on extended runs. First-pass quality on visual and interactive projects also improved: given a product idea, Grok 4.6 establishes structure and visual language in a single pass more reliably than 4.5 did.
Benchmark Numbers
Artificial Analysis Intelligence Index: 61 (up from 56 on Grok 4.5, matching GPT-5.6 Sol Max) DeepSWE v1.1: 65.9% (up from 54.0%) Terminal-Bench v3.0: 26.0% (up from 15.7%) GDPVal-AA v2: 1753 (up from 1526 — leading GPT-5.6 Sol at 1728) AA-Briefcase: leads the comparison table Harvey LAB: leads the comparison table CursorBench v3.2: 69.9% (vs 66.x on Grok 4.5) Where GPT-5.6 Sol still leads: DeepSWE and Terminal-Bench by clear margins in xAI's own table. Fable 5 Max leads on the overall Intelligence Index at 62 and on six task rows. That's honest — this isn't a claim to be best at everything. It is a clear and well-documented step forward from 4.5, and the first time a Grok model has matched Sol on the headline composite index.
The xhigh Reasoning Level
Grok 4.6 adds a fourth reasoning setting above the existing high. xhigh burns more tokens for harder problems — the right move for difficult architecture questions, debugging stubborn failures, or tasks where a single correct answer matters more than throughput. For most coding tasks, high remains the default. xhigh is the setting to reach for when a task has stumped other models or when you want the most careful analysis the model can produce. This matches the effort-control patterns already available in GPT-5.6 and Claude Opus 5 — model choice and reasoning depth can now be set independently on Grok as well.
Pricing — The 200K Caveat
Standard rate: $2 input / $0.50 cached input / $6 output per million tokens. This is unchanged from Grok 4.5 and is competitive against every comparable frontier model. One detail worth knowing: xAI applies a long-context surcharge once a prompt crosses 200K tokens. At that point the rate doubles to $4 input / $1 cached input / $12 output, and the higher rate applies to all tokens in that request — not just the portion above the threshold. For most tasks the standard rate applies. For very long-context work — filling 300K+ tokens with a large codebase or document set — factor in that jump. On Multos, Auto Router manages context length automatically and routes to Grok 4.6 where the task fits comfortably within the short-context band. If you're selecting manually, keep the 200K band in mind for unusually large inputs.
When to Use Grok 4.6 on Multos
Grok 4.6 is the right pick for long-running agentic tasks where the model needs to sustain work across many tool calls — building a feature from scratch, debugging a multi-file issue end-to-end, or researching an unfamiliar domain and producing a structured output. Its improved self-verification means fewer incomplete tool chains on longer jobs. It also earns a place on first-pass visual and interactive projects. The model's ability to establish visual structure and component layout in a single generation pass reduces the iteration cost on frontend work. For quick single-turn tasks or high-volume batch work, DeepSeek V4 Flash and MiniMax M3 remain more efficient. Grok 4.6 is the model to reach for when the task is complex enough to warrant frontier-level sustained reasoning — and when you want that at $2/$6 pricing rather than the $5/$30 of GPT-5.6 Sol.
Available Now
Grok 4.6 is live on Multos for all Starter plans and above. Select it from the model dropdown under xAI, or use xAI Grok OAuth via your X Premium+ subscription if you prefer billing through xAI directly. Auto mode will route to Grok 4.6 where appropriate based on task complexity and cost. The multiplier sits in the mid tier — your token pool on Starter and above gives you meaningful access to it.
Ready to Get Started?
Grok 4.6 is live on Multos Starter and above. Select it from the model dropdown or let Auto Router decide.
Start Building FreeRelated Articles
GPT-5.6 Luna Price Cut 80% — Now Available on Free & Lite
OpenAI slashed GPT-5.6 Luna by 80% overnight — from $1/$6 to $0.20/$1.20 per million tokens. We're passing it straight through: Luna is now available on every plan including Free and Lite.
Read more →Claude Opus 5 & Sonnet 5 Now Available: Near-Frontier Intelligence at Every Tier
Anthropic's Claude Opus 5 and Claude Sonnet 5 are now live on Multos. Opus 5 delivers near-Fable-5 intelligence at half the price with 79.2% on SWE-bench Pro, while Sonnet 5 brings Opus-class agentic capabilities to Starter-tier users at $3/$15 per million tokens.
Read more →Kimi K3 Now Available: The World's Largest Open-Weight Model on Multos
Moonshot AI's 2.8-trillion-parameter Kimi K3 is now live on Multos. A frontier-class open model with a 1M token context window, ranked #1 on Frontend Code Arena and #4 overall on Artificial Analysis Intelligence Index.
Read more →