RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
RoboSPA is a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in Vision-Language-Action (VLA) models. It focuses on Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants. The dataset contains 527K trajectories across multiple embodiments and diverse scenes. Experiments on representative VLA models show current systems struggle with complex spatial relations, precise low-level execution, and memory-intensive planning.
RoboSPA introduces a benchmark to assess VLA models beyond simple scenes and short-horizon tasks. It evaluates fine-grained spatial reasoning and long-horizon procedural planning across 280 task variants derived from 56 base tasks in 10 categories. The dataset includes 527K trajectories from multiple embodiments. Initial results indicate existing VLA models have significant limitations in complex spatial relations, precise execution, and memory-intensive planning.
The benchmark's diagnostic metrics go beyond binary success rate, enabling detailed analysis of failure modes. The five difficulty levels per task systematically increase spatial ambiguity and procedural complexity, providing a structured way to measure model degradation. The inclusion of multiple embodiments tests generalization across robot platforms.
RoboSPA addresses a gap in evaluating VLA models for real-world robotic manipulation, where tasks often require long-horizon planning and fine spatial reasoning. The dataset's scale and diversity could become a standard for benchmarking embodied AI systems, influencing research and development priorities in robotics.
RoboSPA provides a valuable resource for companies developing robotic manipulation systems, enabling systematic evaluation and comparison of VLA models. It may reduce development time by highlighting specific failure modes and guiding targeted improvements.
Future work may involve improving VLA models to handle the identified weaknesses, such as memory-intensive planning and precise low-level execution. The benchmark could be extended with more tasks, embodiments, or real-world data. Adoption by the research community may lead to new model architectures or training methods.