Event date · · ScaleWoB

ScaleWoB: Thousand-Level Verifiable Tasks Show Mobile GUI Agent Long-Task Success Rate Only 17.82%

FACT STATEMENT

Submitted on May 24, 2026, ScaleWoB synthesizes 100+ interaction environments and 1,000+ verifiable tasks; five mobile GUI agents average 27.92% success rate, dropping to 17.82% for long tasks, while humans achieve 92.08%.

What happened

Real applications are difficult to reset and construct rewards, limiting the scale of GUI agent training and evaluation. ScaleWoB uses high-fidelity web simulations without backend to emulate mobile, desktop, and in-vehicle interfaces, reducing environment costs and exposing current capability gaps.

Technical significance

The framework uses a coding agent to generate cross-platform interaction environments and verifiable rewards; environments are accessed via URL with near-zero configuration. The public benchmark includes 120 hard tasks across 63 simulated mobile apps, and validates that rankings from synthetic environments transfer to real application samples.

Industry impact

GUI agent competition will increasingly depend on scalable environments, state reset, and result verification, rather than manually recording a few demos; synthetic environments may become training data and regression infrastructure.

Decision value

Teams can first build resettable digital twins for core workflows, use long-task success rate and recovery rate to screen models, then proceed to low-permission shadow testing on real websites.

What to watch

Need to verify long-term distributional differences between synthetic and real interfaces, visual details, login permissions, and prompt injection, and prevent agents from overfitting to generator patterns.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.