OmegaUse-OfficeVal · Jul 29, 2026

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

OmegaUse-OfficeVal is a benchmark of 100 long-horizon office-suite tasks, averaging 2.32 human hours each, with task-level economic grounding via human labor time and task price proxy. Frontier LLMs are cheaper and faster than humans but have not yet matched human deliverable quality. Code and dataset are open-sourced.

What happened

Researchers introduced OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on complex office-suite workflows. The benchmark includes 100 tasks derived from practitioner requests, each paired with economic signals (human labor time and price proxy) to enable cost comparisons. Evaluations show that while frontier LLMs are significantly cheaper and faster than human workers, their output quality still falls short of human-level deliverables.

Technical significance

The benchmark uses code-based verifiers built from fine-grained rubrics to ensure stable evaluation. Economic grounding is achieved by attaching human labor time and task price proxies to each task, enabling direct cost and value-weighted comparisons between human and LLM performance.

Industry impact

This benchmark highlights a gap between LLM speed/cost advantages and quality in real-world office tasks, suggesting that current agents are not yet reliable for autonomous deployment in enterprise settings without human oversight.

What to watch

Next signals to watch include whether subsequent LLM versions close the quality gap on OmegaUse-OfficeVal, and whether enterprises adopt hybrid human-AI workflows based on such economic benchmarks.

Decision value

The benchmark provides a framework for enterprises to quantify the trade-off between cost savings from LLM automation and potential quality degradation, aiding in ROI analysis for AI adoption in office productivity.

Evidence