Event date · · WorldBench

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

FACT STATEMENT

WorldBench is a multilingual benchmark of persona-grounded everyday workflows with 1,600 tasks across seven languages and eight cultures. It introduces Constrained Task Success (CTS) combining natural language instructions and testbeds. Frontier models reach only 49.2% CTS, showing gaps between correctness and environment preservation.

What happened

WorldBench is a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions. It comprises 1,600 tasks across seven languages and eight cultures, filtered and refined through feedback from human annotators with language- and culture-specific expertise. Evaluation extends previous metrics and introduces Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and other complementary metrics through deterministic and LLM-as-a-Judge evaluations. Experiments show frontier models reach only 49.2% CTS, with all models demonstrating large gaps between correctness and environment preservation, indicating current agents remain brittle in multilingual, agentic scenarios, especially for long-horizon tasks and under state-preservation constraints.

Technical significance

The benchmark's CTS metric integrates deterministic testbeds with LLM-as-a-Judge, enabling measurement of both task completion and environment preservation. The 49.2% CTS for frontier models highlights a significant capability gap in state preservation and multilingual long-horizon reasoning, suggesting current architectures struggle with maintaining context across culturally diverse workflows.

Industry impact

This benchmark exposes a critical weakness in agentic AI for global deployment: performance degrades in non-English and culturally specific scenarios. Enterprises targeting multilingual markets may need to invest in localization and state-management improvements before relying on autonomous agents for complex workflows.

Decision value

For businesses, WorldBench provides a measurable way to assess agent readiness for international operations. The low CTS score signals that current frontier models are not yet reliable for autonomous, multi-step tasks in diverse linguistic and cultural contexts, which may delay enterprise adoption and increase demand for human-in-the-loop systems.

What to watch

Expect follow-up research to focus on improving state preservation and cross-lingual transfer in agent architectures. Adoption of WorldBench as a standard evaluation could drive model developers to prioritize robustness in multilingual, long-horizon tasks, potentially leading to new training paradigms or memory mechanisms.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.