ScaleWoB: Thousand-Level Verifiable Tasks Show Mobile GUI Agent Long-Task Success Rate Only 17.82%
Submitted on May 24, 2026, ScaleWoB synthesizes 100+ interaction environments and 1,000+ verifiable tasks; five mobile GUI agents average 27.92% success rate, dropping to 17.82% for long tasks, while humans achieve 92.08%.
Real applications are difficult to reset and construct rewards, limiting the scale of GUI agent training and evaluation. ScaleWoB uses high-fidelity web simulations without backend to emulate mobile, desktop, and in-vehicle interfaces, reducing environment costs and exposing current capability gaps.
The framework uses a coding agent to generate cross-platform interaction environments and verifiable rewards; environments are accessed via URL with near-zero configuration. The public benchmark includes 120 hard tasks across 63 simulated mobile apps, and validates that rankings from synthetic environments transfer to real application samples.
GUI agent competition will increasingly depend on scalable environments, state reset, and result verification, rather than manually recording a few demos; synthetic environments may become training data and regression infrastructure.
Teams can first build resettable digital twins for core workflows, use long-task success rate and recovery rate to screen models, then proceed to low-permission shadow testing on real websites.
Need to verify long-term distributional differences between synthetic and real interfaces, visual details, login permissions, and prompt injection, and prevent agents from overfitting to generator patterns.