Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis
A paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp framework that learns a generalizable base policy for grasp synthesis across different robotic hands, using explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to composable foundation-model modules providing spatial, cognitive, and temporal priors.
The paper introduces AdaRoboVLG, a framework for adaptive vision-language grasping that separates generalizable grasp synthesis from task-specific understanding. The base policy generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, enabling efficient learning and cross-hand generalization. Task-dependent understanding is handled by specialized foundation-model modules that provide composable priors (spatial, cognitive, temporal) integrated into grasp synthesis without retraining the base policy. Simulation and real-world experiments show the framework addresses three representative grasping challenges and supports functional grasping when priors operate jointly, with performance comparable to state-of-the-art methods.
The key technical contribution is decoupling grasp synthesis from task understanding: a generalizable base policy uses explicit kinematic mapping and force-closure stability estimation to generate physically feasible grasps across different robotic hands, while composable foundation-model priors inject task context without retraining. This modular design enables efficient learning and cross-hand generalization, and the priors can be combined to enable functional grasping.
This approach could reduce the need for hand-specific grasp policy retraining in robotic manipulation, potentially lowering deployment costs for multi-hand robotic systems. The use of composable foundation-model priors suggests a path toward more flexible, task-adaptive robotic grasping in industrial and service robotics.
The framework's cross-hand generalization and task adaptability could reduce engineering effort and time-to-market for robotic grasping solutions across different hardware, offering potential cost savings for robotics companies and integrators.
Next signals to watch include real-world deployment results on diverse robotic hands, integration with commercial robot platforms, and extensions to more complex manipulation tasks beyond grasping. The paper's emphasis on composable priors may influence future research on modular robot learning architectures.