Grok 4.3 vs Grok 4.2: what changed in price, speed, and agentic scores
xAI shipped Grok 4.3 on April 30, 2026. Independent analysts at Artificial Analysis report an Intelligence Index score of 53 for Grok 4.3, improved agentic benchmark results versus Grok 4.20 0309 v2, and a lower dollar cost to run the full Artificial Analysis Intelligence Index suite at about $395 compared with the prior Grok generation under the same methodology. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
The headline is not “a new version number.” It is a shift in cost-per-intelligence on Artificial Analysis’ composite leaderboard: higher measured capability alongside cheaper full-suite evaluation costs driven by lower per-token pricing, even when the model consumes more output tokens than Grok 4.20 0309 v2 on the same benchmark battery. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
Primary sources: Artificial Analysis’ unrolled thread summary and the Grok 4.3 model card on Artificial Analysis. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
What shipped
Artificial Analysis positions Grok 4.3 as a proprietary reasoning model that moves xAI up the Intelligence Index while improving benchmark economics. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
- Intelligence Index: Grok 4.3 scores 53 on the Artificial Analysis Intelligence Index, placing it just above Muse Spark and Claude Sonnet 4.6 and about four points ahead of the latest Grok 4.20 in Artificial Analysis’ ranking narrative. (Source: Artificial Analysis thread)
- API sticker prices: Artificial Analysis lists Grok 4.3 at $1.25 per million input tokens and $2.50 per million output tokens, with a cache hit input rate of $0.20 on the model page snapshot used for this article. (Source: Artificial Analysis Grok 4.3)
- Full-suite evaluation cost: Artificial Analysis reports about $395 to run the Intelligence Index for Grok 4.3 in the thread narrative, and $395.17 on the model page, framed as roughly 20% lower than Grok 4.20 0309 v2 for the same suite. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
- Price cuts versus Grok 4.20: The thread cites 37.5% lower input token prices and 58.3% lower output token prices as drivers, alongside an older headline-style claim of about 40% lower input and 60% lower output versus Grok 4.20. (Source: Artificial Analysis thread)
- Throughput: The Grok 4.3 model page lists 189.9 output tokens per second on xAI’s API in the captured snapshot, which Artificial Analysis ranks highly on speed among evaluated models. (Source: Artificial Analysis Grok 4.3)
- Verbosity: Grok 4.3 uses about 44% more output tokens than Grok 4.20 0309 v2 to complete the Intelligence Index, while the model page records 88M output tokens for the Intelligence Index run and still frames Grok 4.3 as comparatively less verbose than some other leading models in Artificial Analysis’ narrative. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
- Modalities and context: The model page lists text and image input, text output, a 1M token context window, and classifies Grok 4.3 as a reasoning model. (Source: Artificial Analysis Grok 4.3)
Grok 4.3 vs Grok 4.2: the head-to-head
Every figure below is Artificial Analysis' own reporting for Grok 4.3 against Grok 4.20 0309 v2, the Grok 4.2-generation checkpoint it benchmarked as the predecessor. The right-hand column states the change in the direction Artificial Analysis states it, so nothing is back-computed. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
| Dimension | Grok 4.3 | Change vs Grok 4.20 0309 v2 |
|---|---|---|
| Artificial Analysis Intelligence Index | 53 | about 4 points higher |
| GDPval-AA Elo (agentic work tasks) | 1500 | up 321 points from 1179 |
| Tau-squared-Bench Telecom | 98% | up 5 points |
| IFBench (instruction following) | 81% | unchanged, carried forward |
| AA-Omniscience Accuracy | not published as an absolute | up 8 points |
| AA-Omniscience Non-Hallucination Rate | not published as an absolute | down 8 points |
| Input price per 1M tokens | $1.25 | 37.5% lower |
| Output price per 1M tokens | $2.50 | 58.3% lower |
| Cache-hit input per 1M tokens | $0.20 | not compared in this snapshot |
| Intelligence Index suite cost | about $395 | about 20% lower |
| Output tokens consumed by the suite | 88M | about 44% more |
| Context window | 1M tokens | not compared in this snapshot |
The three rows that decide most migrations are output price, GDPval-AA, and the non-hallucination line. Output price fell by more than input price, which favors long-generation and agent-loop workloads; GDPval-AA moved the most of any benchmark on the card; and non-hallucination moved the wrong way, which is the one row that can disqualify Grok 4.3 for policy-grounded answering. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
Read the two price-cut figures carefully. The thread cites 37.5% lower input and 58.3% lower output, alongside an older headline-style claim of roughly 40% lower input and 60% lower output. Those are two different roundings of the same cut, not two separate cuts. (Source: Artificial Analysis thread)
Status check: Grok 4.3 is now marked deprecated
Operator note (first-hand): We re-fetched https://artificialanalysis.ai/models/grok-4-3 on 2026-07-27. The page now carries the line "This model is deprecated. We only continue performance benchmarking for the default 10k input token workload," and states that a newer model, Grok 4.5 (high), has launched. Published pricing was unchanged from our May capture at $1.25 input, $2.50 output, and $0.20 cache-hit input per 1M tokens, and the context window still reads 1M tokens. (Source: Artificial Analysis Grok 4.3)
The benchmark figures on that page have moved, and the reason is stated on the page itself: Artificial Analysis now reports only the default 10k input token workload for the deprecated Grok 4.3 (high) variant. Under that narrower harness the page shows an Intelligence Index of 38, output speed of 104.0 tokens per second, an Intelligence Index run cost of $303.34, and 83M output tokens for the run. Those are not corrections to the May figures; they are a different measurement configuration for the same model. (Source: Artificial Analysis Grok 4.3)
What that means in practice: treat the May 2026 numbers in this article as the full-suite snapshot taken at release, and treat the July figures as the reduced-workload continuation. If you are choosing a Grok tier today, price Grok 4.5 against your own evaluation harness rather than inheriting either snapshot. (Inference: reading of Artificial Analysis' stated benchmarking-scope change; no claim about Grok 4.5's scores is made here.)
Practitioner payoff: score up, benchmark bill down
Teams that route traffic by “leaderboard tier” and API price should treat Artificial Analysis’ Intelligence Index as one composite signal, not a replacement for task-specific evals. Still, the Grok 4.3 story is unusually concrete on economics: Artificial Analysis explicitly ties the Intelligence Index run cost to combined token usage and per-token pricing, and reports a lower suite cost for Grok 4.3 than Grok 4.20 0309 v2 despite higher output-token usage on that suite. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
Practitioner payoff: If your internal workloads resemble long outputs and agent-style turns, output-token price and verbosity dominate bills more than input-token price. Grok 4.3’s published output price is $2.50 per million output tokens on Artificial Analysis’ card, and the thread emphasizes that output-token volume rose versus Grok 4.20 0309 v2 even as total suite cost fell. That pattern rewards teams that measure dollars per successful task, not dollars per million tokens in isolation. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
Why this matters: A lower Intelligence Index run cost is not the same as “cheap in production,” but it is a meaningful signal that xAI is pushing Grok 4.3 toward a more favorable spot on Artificial Analysis’ intelligence-versus-cost charts, which many buyers use as a first-pass filter before deeper evaluations. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
Operator note (first-hand): On 2026-05-02 we fetched https://artificialanalysis.ai/models/grok-4-3 over HTTPS and captured the published Intelligence Index score (53), pricing lines ($1.25 input, $2.50 output), output speed (189.9 tokens per second), Intelligence Index output-token total (88M), and suite cost total ($395.17) directly from the live page content returned to the client. (Source: Artificial Analysis Grok 4.3)
Agentic lifts: GDPval-AA, instruction following, and support simulations
Artificial Analysis highlights Grok 4.3’s largest single benchmark jump on GDPval-AA, its agentic evaluation focused on real-world tasks. Grok 4.3 posts an Elo of 1500, up 321 points from 1179 for Grok 4.20 0309 v2. Artificial Analysis says Grok 4.3 surpasses Gemini 3.1 Pro Preview, Muse Spark, GPT-5.4 mini (xhigh), and Kimi K2.5 on that benchmark snapshot, while still trailing GPT-5.5 (xhigh) by 276 Elo points with an expected win rate of about 17% head-to-head under a standard Elo framing. (Source: Artificial Analysis thread)
Practitioner payoff: If your product roadmap looks like “agents that complete multi-step workflows in messy domains,” GDPval-AA is closer to that risk surface than a pure coding leaderboard. The magnitude of the jump matters: 321 Elo points is a headline-grade move, even if GPT-5.5 (xhigh) remains the leader on Artificial Analysis’ snapshot. (Source: Artificial Analysis thread)
On instruction following and customer-support style simulations, Artificial Analysis reports that Grok 4.3 gains five points on 𝜏²-Bench Telecom to 98%, described as in line with GLM-5.1, and maintains an 81% IFBench score carried forward from Grok 4.20 0309 v2. Those lines support the thread’s theme that Grok 4.3 is intentionally competitive on agentic customer-support scenarios, not only on abstract reasoning scores. (Source: Artificial Analysis thread)
Decision rule for teams: Treat telecom-style tool simulations as a directional signal for regulated or procedure-heavy support flows, then validate on your own transcripts, ticket taxonomy, and tool contracts. Benchmark leaders can still fail where your tools differ or where compliance constraints narrow allowable actions. (Inference: common deployment practice when benchmarks approximate but do not equal production.)
Knowledge stack tradeoff: AA-Omniscience accuracy versus non-hallucination rate
Artificial Analysis also reports a mixed picture on AA-Omniscience, its knowledge-and-hallucination framing. Grok 4.3 gains eight points on AA-Omniscience Accuracy, but loses eight points on AA-Omniscience Non-Hallucination Rate versus Grok 4.20 0309 v2. On Non-Hallucination Rate, Grok 4.20 0309 v2 still leads in Artificial Analysis’ snapshot, followed by MiMo-V2.5-Pro, with Grok 4.3 positioned on that leaderboard rather than at the top of that specific column. (Source: Artificial Analysis thread)
Why this matters: Teams evaluating “accuracy-first” assistants versus “refuse-when-unsure” assistants should separate those objectives. A gain on accuracy without a gain on non-hallucination suggests different failure modes: more correct answers when the model commits, but not necessarily safer abstention behavior. (Inference: interpretation of reported metric directions from Artificial Analysis.)
Defensive focus: If you ship customer-facing answers grounded in policy documents, run red-team prompts that reward abstention, measure hallucination-style failures on your own corpus, and do not assume leaderboard non-hallucination rankings transfer across domains. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
The practical read is not “Grok 4.3 wins every column.” It is that xAI shipped a model that improves Artificial Analysis’ headline intelligence score and several agentic tracks while cutting the analyst suite’s run cost, with explicit tradeoffs visible on omniscience-style metrics. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
Context: benchmark suites increasingly double as pricing reviewers
Artificial Analysis has spent years turning “model comparisons” into repeatable methodology around blended price ratios, cache-aware pricing, token-use measurements, and composite indices. Grok 4.3 is another data point in that meta-story: vendors compete not only on capability slides but on how expensive their models are to run through the same public evaluation harness. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
If you want adjacent framing on how frontier releases interact with operator economics, our GPT-5.5 agentic shift coverage walks through how OpenAI positioned GPT-5.5 across coding and knowledge work, which pairs well with reading GDPval-AA movements as part of a broader agentic trendline. (Inference: editorial pointer; no benchmark equivalence implied.)
Adoption notes
Decision rules for teams:
- Pin your evaluation harness before you pin the model. If Grok 4.3’s strengths on Artificial Analysis are agentic tracks and instruction-following simulations, mirror those tasks with your own tools and data before switching production routes. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
- Recompute total cost with your token mix. Artificial Analysis’ suite cost combines usage and published pricing; your application may emit shorter or longer outputs than the Intelligence Index harness, which changes where input versus output pricing bites hardest. (Sources: Artificial Analysis thread, Artificial Analysis Grok 4.3)
- Treat omniscience metrics as a policy question. If non-hallucination rate moved in the wrong direction for your risk appetite, add retrieval grounding, citation requirements, and escalation paths independent of the headline Intelligence Index score. (Source: Artificial Analysis thread)
- Compare against Sonnet-class alternatives on real workflows. Artificial Analysis places Grok 4.3 near Claude Sonnet 4.6 on the Intelligence Index narrative in the thread; the right choice still depends on latency, compliance posture, and toolchain fit. (Source: Artificial Analysis thread)
Frequently asked questions
Is Grok 4.2 better?
Not on Artificial Analysis' composite. Grok 4.3 scores 53 on the Intelligence Index, about four points above the Grok 4.20 generation, and gains 321 GDPval-AA Elo points. Grok 4.20 0309 v2 does keep the lead on AA-Omniscience Non-Hallucination Rate, where Grok 4.3 dropped eight points. (Source: Artificial Analysis thread)
What is the best Grok model right now?
Not Grok 4.3. As of 2026-07-27 the Artificial Analysis model page marks Grok 4.3 deprecated and points to a newer release, Grok 4.5 (high). Grok 4.3 remains the right reference point for understanding the price and agentic-benchmark shift, but new deployments should evaluate the current tier. (Source: Artificial Analysis Grok 4.3)
What can Grok 4.3 do?
Artificial Analysis classifies it as a reasoning model taking text and image input and returning text, with a 1M-token context window. Its strongest reported tracks are agentic: 1500 GDPval-AA Elo, 98% on Tau-squared-Bench Telecom, and 81% IFBench for instruction following. (Source: Artificial Analysis Grok 4.3)
How much does Grok 4.2 cost?
Artificial Analysis publishes the delta, not the Grok 4.20 sticker price: Grok 4.3 is 37.5% cheaper on input and 58.3% cheaper on output. Inference: applying those cuts to Grok 4.3's $1.25 input and $2.50 output implies roughly $2.00 and $6.00 per 1M tokens for Grok 4.20. (Source: Artificial Analysis thread)
How much does Grok 4.3 cost per million tokens?
Artificial Analysis lists $1.25 per 1M input tokens, $2.50 per 1M output tokens, and $0.20 per 1M cache-hit input tokens. Those figures were unchanged when we re-checked the page on 2026-07-27. Running the full Intelligence Index suite cost about $395 at release. (Source: Artificial Analysis Grok 4.3)
Related coverage
- GPT-5.5 Arrives: The Agentic Shift in Coding, Research, and Knowledge Work - how OpenAI framed GPT-5.5 across agentic workloads that overlap GDPval-style narratives.
- Claude API Pricing: Haiku, Sonnet 4.6, and Opus 4.8 for Agent Builders - the per-token comparison point when you price Grok 4.3 against Anthropic tiers.
- Claude Opus 4.7 is GA: the migration checklist for agentic coding - migration discipline when upgrading flagship tiers used behind coding agents.
- DeepSeek V4 pricing turns 1M-token context into an operator choice - a contrasting open-weights story where API sticker price and cache mechanics dominate routing decisions.
References
- Artificial Analysis Grok 4.3 - https://artificialanalysis.ai/models/grok-4-3
- Artificial Analysis thread - https://threadreaderapp.com/thread/2049987001655714250.html



