OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
OSReward is a benchmark introduced to evaluate vision-language model (VLM) judges on computer-using agent (CUA) trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, labeled with ground-truth verdicts through multi-stage human annotation. The benchmark includes OSReward-Hard (hard cases) and OSReward-Multi (efficiency and alignment scoring). Evaluation finds even state-of-the-art VLMs fall short of an ideal judge, exhibiting systematic leniency bias.
Researchers introduced OSReward, a benchmark for evaluating VLM judges of CUA trajectories, with subsets for hard cases and multi-dimensional scoring. The most comprehensive evaluation to date reveals that current state-of-the-art VLMs are not reliable enough, showing a systematic leniency bias.
The benchmark exposes a systematic leniency bias in VLM judges, indicating that current models tend to over-approve trajectories. This suggests that reward model training for CUAs may need debiasing techniques or more nuanced scoring rubrics. The multi-stage human annotation process sets a high standard for ground-truth labeling.
As CUAs become more prevalent, reliable automated evaluation is critical for scaling data curation and reinforcement learning. The identified unreliability of VLM judges could slow deployment of autonomous agents in enterprise and consumer applications, as trust in automated verification is essential.
Reliable automated evaluation of CUA trajectories can reduce the cost and time of human annotation, enabling faster iteration and safer deployment of computer-using agents in business process automation, software testing, and digital assistants.
Next signals include research into debiasing VLM judges, development of more robust reward models, and potential adoption of OSReward as a standard benchmark in the CUA community. Improvements in judge reliability could accelerate CUA development and deployment.