SpaceXAI has released Grok 4.6 with a clear priority: agents that stay with complex work across many steps. The model researches topics, analyzes information, navigates codebases and turns broad product ideas into working applications.

Trained to finish the job

The agent focus starts in training. Grok 4.6 went through a longer supplemental training run than Grok 4.5, built on curated model generated data for reasoning and advanced technical concepts, high quality engineering data and an improved optimizer. SpaceXAI then used Grok 4.5 to regenerate the supervised fine tuning trajectories across reasoning efforts, agent harnesses and domains such as STEM, software engineering and knowledge work. Model based checks filtered out problematic traces before the next stage.

Reinforcement learning pushed the model into agentic environments covering general coding, knowledge work, kernel optimization, web development and computer aided design. The reward signal favored completing the larger task over producing a plausible block of code. That choice defines what an agent needs. A model that writes correct code on the first try is useful. A model that catches its own mistakes, fixes them and continues toward the original goal is what makes an agent dependable over hours of work.

Agents that check their own work

On longer trajectories SpaceXAI observed more self testing and verification. The model pauses to check whether its work holds up before moving on, and it corrects mistakes while still in the loop. These are company observations, so independent testing will need to confirm how consistently this behavior shows up in production.

The practical result appears in first versions. Given a concrete product idea, Grok 4.6 can research an unfamiliar domain, structure the application, implement the core interactions and set a visual language in one pass. It then refines the result through several rounds of feedback. For interactive and visual projects the first attempt lands stronger than what Grok 4.5 typically produced, which makes the workflow of starting from something substantial and iterating in the loop realistic.

Benchmark results for agentic work

The numbers tell two stories at once. The improvement over Grok 4.5 is large and consistent. The lead over rival frontier models appears on some tests and stays out of reach on others.

  • Artificial Analysis Intelligence Index. Grok 4.6 scores 61, five points above Grok 4.5 and tied with GPT-5.6 Sol Max. Claude Opus 5 leads at 63 and Fable 5 sits at 62.
  • GDPval-AA v2. An Elo of 1753 on real world knowledge work, behind only Claude Opus 5 and statistically level with Fable 5 and Qwen3.8 Max.
  • CursorBench v3.2. A score of 69.9%, up from 66.7% for Grok 4.5. Fable 5 Max leads at 70.5%.
  • DeepSWE v1.1. A sharp rise from 54% to 65.9%, with Sol Max still ahead at 73%.
  • FrontierCode v1.1 Extended. Grok reaches 61.3%, edging Sol at 60.6% while Fable holds 63.6%.
  • APEX-Agents. A jump of 10.4 points to 57.5%, narrowly ahead of Sol Max at 56.7% and behind Fable at 59.2%.
  • Terminal-Bench v3.0. The biggest remaining gap. Grok improves from 15.7% to 26%, while Sol and Fable score near 34%.
  • τ³-Banking. A score of 50.7% on multi turn customer service with tool use, placing Grok in the top two alongside Qwen3.8 Max.
  • Harvey LAB. Grok reaches 15.8%, compared with 11.3% for Fable 5 Max and 2.5% for Sol Max.

One caveat applies to every figure above. SpaceXAI notes that competitor scores come from published system cards and leaderboards, so treat the comparison as indicative for your own evaluation.

Efficiency on long horizon work

On AA-Briefcase, a private benchmark of extended agentic knowledge work, Grok 4.6 debuts at an Elo of 1577. That puts it at Fable 5 tier and ahead of Sol Max at 1502, with balanced results across rubric grading, presentation quality and analytical quality.

The efficiency profile stands out as much as the score. Grok 4.6 resolved tasks in about 53 turns and roughly 0.5 billion input tokens on average. Claude Opus 5 Max needed around 103 turns and 2 billion input tokens for the same evaluation. Context accumulates fast in long running agent work, so a model that reaches a comparable answer in half the turns and a quarter of the tokens gains a cost advantage. Artificial Analysis places Grok 4.6 on its intelligence versus cost Pareto frontier at $0.84 per task. Cheaper per task options exist, such as GPT-5.6 Luna, GLM-5.2 and Muse Spark 1.2, and each of them scores lower on the Intelligence Index.

What running Grok 4.6 agents costs

Headline pricing stays flat versus Grok 4.5 at $2 per million input tokens and $6 per million output tokens, with cache hits at $0.50. A fast variant costs twice as much. Those rates sit more than 60% below the cited frontier prices of Claude Opus 5 at $5 and $25 and GPT-5.6 Sol at $5 and $30.

An important detail hides in the context window. Grok 4.6 supports 500,000 tokens, and the $2 and $6 rates apply to prompts below 200,000 tokens. Once a single request crosses that line, every token in it bills at $4, $1 for cached input and $12. Keep that threshold in mind when you estimate the total cost of agents that carry heavy context.

The deeper point is that token price and task cost are different metrics. A cheap model that repeatedly calls tools, rereads files and restarts tasks can end up expensive. A pricier model that finishes in fewer steps can cost less overall. The figure worth tracking is the cost of a completed workflow.

Where you can run Grok 4.6 agents

Grok 4.6 is live in Cursor and in Grok Build, the counterpart to Claude Code and Codex, which ships with the $30 per month SuperGrok plan. SpaceXAI doubles the included usage in both tools during the first week. The API, OpenRouter, Vercel and Cloudflare round out the access points.

The API specifications support agent workloads directly: text and image inputs with text output, function calling, structured outputs and reasoning, with rate limits of 150 requests per second and 50 million tokens per minute across us-east-1 and us-west-2. The release arrived one day after Grok Bot, a system for assigning ongoing tasks to persistent agents. Together they show SpaceXAI building the infrastructure around agents, with the model as one layer of a larger stack.

What to weigh before you commit

Three factors deserve attention alongside the benchmark tables. Terminal heavy work still favors Sol and Fable, so match the model to your workload. Benchmarks reflect specific harnesses, and CursorBench measures performance inside an environment owned by a company SpaceX acquired, which makes your own harness, prompts, retries and caching the deciding variables. Finally, the Grok brand carries a public history of safety controversies and open regulatory investigations in the United Kingdom and the European Union. Teams in regulated or customer facing settings will weigh that vendor track record next to any capability gain.

The real test is your invoice

Grok 4.6 makes a specific promise: frontier level agent work at mid tier prices, with the turn and token efficiency to back it up. The generational leap over Grok 4.5 is solid and measurable. The open question is whether the efficiency seen in controlled evaluations carries into your production workloads. If it does, the most convincing benchmark for Grok 4.6 agents will be the inference bill at the end of the month.