LingBot-VLA 2.0 Open Source: Domestic Embodied Model Enters Cross-Embodiment Scaling Stage
Ant Group's Robbyant releases LingBot-VLA 2.0 technical report, pretrained weights, and code; official disclosure states training data covers 20 robot configurations and approximately 60,000 hours of robot/first-person video.
This release maps single-arm, dual-arm, semi-humanoid, and humanoid robots to a unified action space, and opens up 6B weights along with training and deployment code. Individual benchmarks are only part of the picture; results are still primarily based on the team's own tests and require independent reproduction.
Unified 55-dimensional action representation, sparse MoE action expert, and depth/video teacher distillation point to a clear direction: embodied foundation models are simultaneously scaling up data, embodiment, and temporal prediction dimensions.
Domestic embodied intelligence competition is shifting from demos and single-robot policies to reusable foundation models and open-source ecosystems; this redraws the boundaries between model layers, data layers, and robot hardware manufacturers.
CEOs and investment leads should prioritize 'cross-embodiment transferability, open-weight availability, and real-world deployment costs' as core due diligence criteria for embodied projects, rather than relying solely on team-reported simulation success rates.
Key observations include third-party reproduction success rates on new robots, cross-embodiment fine-tuning costs, real long-horizon task completion rates, and whether open-source weights foster a developer ecosystem.