Grok 4 Released: Native Tool Use and Large-Scale Reinforcement Learning Become the Main Line
xAI released Grok 4 and Grok 4 Heavy in July 2025, providing access via the Grok product and API.
Grok 4 integrates reasoning training, code interpreter, web browsing, and X search into a unified model capability. xAI's competitive focus shifts from model chat experience to a research and execution system with callable tools.
Officially, Grok 4 scales reinforcement learning training on Colossus and uses native tool use to handle cross-domain verifiable tasks; the Heavy version extends result quality through higher compute budget.
Frontier labs begin to treat search, code, and task execution as model training objectives rather than post-release add-ons. Agent reliability becomes a new differentiator.
Enterprise pilots should record the tool call chain, failure points, and human handoffs for each task, avoiding using single-turn reasoning benchmarks as production acceptance.
Observe tool selection accuracy, search source coverage, long-task failure recovery, Heavy mode cost, and third-party reproducible evaluations.