September 21: xAI (SpaceXAI) officially launched Grok 4.7This is the next-generation flagship model following Grok 4.6 on August 12, officially positioned as “a cutting-edge model for programming, agent tasks, and knowledge work.” Its most striking feature is not raw capability—but its pricing: $2 per million input tokens and $6 per million output tokens, unchanged from the previous generation yet only a fraction of GPT-5.6 Sol ($4/$20) and Claude Fable 5.1 ($10/$50). In one sentence: xAI aims to directly capture the most lucrative market—programming and knowledge work—with a “cutting-edge performance + rock-bottom pricing” one-two punch.
Core capability: Specialized in “multi-hour” hard tasks
The biggest change in Grok 4.7 lies not in parameter count (officially undisclosed; industry estimates place it at ~2.1 trillion parameters, roughly a 40% increase over Grok 4.6), but intraining methodologyxAI adopted a larger base model this time and conducted longer reinforcement learning (RL) training, with training data deliberately skewed toward problems requiring “hours to solve.”
This yields three direct capability improvements:
- Stronger self-verificationThe model excels at checking its own outputs, reducing “plausible but incorrect” hallucinations—a critical advantage for extended programming tasks.
- More stable long-context managementWith a 500K-token context window, maintains consistency across ultra-long documents and large codebases.
- Native understanding of Grok Bot harnessxAI specifically trained it to understand its own agent framework, resulting in better performance on conversational tasks and general knowledge work, as well as significantly improved document and PPT generation capabilities.
Benchmark scores: Approaching the state of the art, but not universally dominant
xAI officially released a comparison table of Grok 4.7 versus Grok 4.6, GPT-5.6 Sol,Claude and Fable 5.1. The conclusion is “approaching the state of the art, with wins and losses across categories,” not one-sided dominance:
- Long-duration programming tasks46.3% (Grok 4.6: 40.4%, GPT-5.6 Sol: 41.7%, Fable 5.1: 51.8%) — a substantial improvement over the previous generation, yet still trailing Fable 5.1.
- Terminal-Bench 4.0(Multi-hour agent terminal operations): 38.0%, nearly double Grok 4.6’s 20.3%, representing the strongest generational improvement.
- Engineering evaluation DE71.0%, surpassing Grok 4.6 (65.2%) and approaching GPT-5.6 Sol (72.7%).
- Harvey Legal Agent Benchmark19.6%, substantially outperforming Fable 5.1 (6.7%) and GPT-5.6 Sol (2.5%), its strongest individual metric.
- Clinical reasoning HealthBench56.7%, trailing Fable 5.1 (62.1%) and GPT-5.6 Sol (60.5%).
Additionally, xAI claims Grok 4.7 featuresa new safety guardrail stackscoring 62.4% on the LatchBio biosafety benchmark (leading), and permitting only 3.3% of hazardous dual-use prompts on its own HackerBench v0.3, while rarely blocking legitimate safety research. Safety and compliance are emphasized more than ever compared to the previous generation.
User experience and limitations
Smooth onboarding path: Grok 4.7 is now integrated CursorGrok Build (x.ai/build, free to try), Grok APIand various third-party programming harnesses and model routers (e.g., OpenRouter). For speed, xAI also offers a “fast variant” that doubles output speed—and doubles the price ($4/$12).
Some reality checks are warranted:Chinese-language capability remains a weak spotGrok models have historically excelled at English and STEM tasks; Chinese generation quality lags noticeably behind domestic models (e.g., GLM,Kimi,DeepSeek). Second, it still falls short of first place on certain comprehensive benchmarks (e.g., long-horizon coding, clinical reasoning); if you prioritize “absolute strongest” over “best value,” Fable 5.1 or GPT-5.6 Sol remain superior.
Grok 4.7 vs. peer models—head-to-head comparison
| Model | context | Pricing ($/M Token) | Core Positioning | This Site’s Rating |
|---|---|---|---|---|
| Grok 4.7 | 500K | 2 / 6 | Coding + knowledge work | 8.7 |
| Claude Fable 5.1 | — | 10 / 50 | Enterprise agents | 8.9 |
| Gemini 3.8 Live | — | — | Real-time voice | 8.5 |
| GLM 5.3 FlashX | — | — | High-Speed Multimodal | 8.6 |
| Kimi K3 | — | — | Domestic Inference | 8.5 |
Plotting it on a coordinate system, Grok 4.7 occupies a shrewd position: capabilities approach Tier-1 levels, yet pricing is only a fraction of competitors’. For developers and knowledge workers who need affordability, robustness, and sustained task performance, it may be one of the most balanced value propositions available today.
Overall Score
| 维度 | Score | evaluate |
|---|---|---|
| functional completeness | 8.8 / 10 | All-rounder for coding, agents, and knowledge work—multimodal support remains weak |
| 易用性 | 8.6 / 10 | CursorMultiple access points: Grok Build, API—fast onboarding |
| Cost-effectiveness | 9.0 / 10 | $2/$6 pricing is significantly lower than peers |
| 中文支持 | 7.8 / 10 | Strong in English, weaker than domestic models in Chinese |
| 输出质量 | 8.7 / 10 | State-of-the-art overall, but not top-ranked on all benchmarks |
Overall rating: 8.7/10
Frequently Asked Questions (FAQ)
Is Grok 4.7 free?
You can try Grok 4.7 for free via Grok Build (x.ai/build); Cursorusing Grok API or OpenRouter requires payment—API pricing is $2 per million input tokens and $6 per million output tokens.
What is the context window size of Grok 4.7?
Officially labeled as a 500K-token context window, sufficient to process ultra-long documents or large codebases in a single pass.
Grok 4.7 vs. GPT-5.6 Sol andClaude Fable 5.1—which is better?
Each has strengths: Grok 4.7 leads on Terminal-Bench and legal agent tasks and is the most affordable; GPT-5.6 Sol and Fable 5.1 excel at comprehensive tasks like long-horizon programming and clinical reasoning. Choose Grok 4.7 for value-for-money; choose the latter two for absolute top-tier performance.
Does Grok 4.7 support Chinese? How well does it perform?
It supports Chinese, but its Chinese generation quality lags behind domestic models such as GLMKimi,DeepSeek —its strengths lie in English and STEM tasks.
Want to Discover More Useful AI Tools? Explore Our AI Model Library and Tool Comparison Engineor continue reading:Claude Fable 5.1 Evaluation · GLM 5.3 FlashX benchmarking · All Reviews.

Pingback: Tesla Launches Grok Bot: xAI Agent Enters Vehicles for $300/Month — AI Dash