Technical Report on the CVPR 2026@AdvML Workshop Challenge
The CVPR 2026@AdvML Workshop Challenge released a technical report. The challenge targets adversarial multi-modal attacks on autonomous driving vision-language models (VLAs), based on DriveLM-style multi-view visual question answering. It consists of two phases, with the second phase introducing a hidden black-box model to evaluate transferability. The report describes task design, submission rules, evaluation protocol, and leaderboard results, and analyzes five submitted technical reports.
The CVPR 2026@AdvML Workshop Challenge technical report is released, focusing on adversarial attacks on autonomous driving VLAs, using multi-view images and text perturbations, with two phases to evaluate attack transferability.
Image-side attacks are more favored due to suffix penalty; scene-level multi-view optimization outperforms single-view processing; QA types and graph structures provide useful priors for attacks.
Safety evaluation of autonomous driving VLAs is becoming a research hotspot, with adversarial attack challenges driving robustness research.
Improving the safety of autonomous driving systems and reducing the risk of adversarial attacks has potential value for the commercialization of autonomous driving.
In the future, more standardized attack benchmarks and defense methods for multi-modal VLAs may be seen.