GLM-5.1 Released: Zhipu Advances Agent Goal to 8 Hours of Continuous Execution
Zhipu's official Release Notes recorded the release of GLM-5.1 in April 2026, positioning it as a long-duration task flagship model.
Zhipu has begun measuring model upgrades by continuous execution duration, iterative optimization, and engineering delivery. The 8-hour goal also makes state drift, error accumulation, and safety boundaries key product focuses.
Official documentation lists 200K context, 128K maximum output, function calling, MCP, and multi-mode thinking, and states that the model can continuously execute a single task for up to 8 hours.
Model evaluation is shifting from minute-level benchmarks to hour-level real-world tasks, increasing the importance of Agent runtime, sandboxing, cost control, and process evaluation.
Enterprises should first establish stage checkpoints, budget caps, and human handover before expanding long-duration autonomy; continuous runtime alone is not business value.
Observe the public trajectory of 8-hour tasks, success rate distribution, strategy drift, resource costs, and recoverability after errors.