VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
VBVR-Pro is a closed-loop testbed for native visual reasoning through generation. It includes 300 procedurally generated tasks and verifiable reward scorers. Models trained on VBVR-Pro show transfer to seven external visual reasoning benchmarks including RISE-Video, MME-CoF-Pro, and BabyVision. The work identifies failure modes of VLM-as-a-judge and proposes deterministic, task-specific scorers.
VBVR-Pro introduces a scalable and verifiable suite for native visual reasoning, treating visual generation as the medium of reasoning. It provides 300 procedurally generated tasks and verifiable reward scorers, enabling trainable, verifiable, optimizable, and experimentally controllable visual reasoning. Models trained on VBVR-Pro demonstrate strong transfer to external benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. The study also reveals recurring failure modes in using leading MLLMs as judges and advocates for deterministic, task-specific scorers.
The suite's procedural task generation and deterministic scorers address the lack of scalable training tasks and reliable feedback in native visual reasoning. The reported transfer to seven external benchmarks suggests the learned capabilities generalize beyond the training distribution. The critique of VLM-as-a-judge highlights the need for grounded evaluation in visual reasoning tasks.
This work may accelerate research in visual reasoning by providing a standardized testbed, potentially influencing how future multimodal models are evaluated and trained. The emphasis on verifiable rewards could shift evaluation practices away from subjective model-based judging.
The testbed could lower the barrier to developing and validating visual reasoning models, potentially enabling new applications in areas requiring visual problem-solving, such as robotics, design, and autonomous systems. Standardized evaluation may also facilitate technology transfer and benchmarking in industry.
Observable next signals include adoption of VBVR-Pro in subsequent research, improvements in native visual reasoning benchmarks, and development of more robust verifiable reward mechanisms. The suite may become a common benchmark for comparing generative visual reasoning approaches.