Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation
A research paper introduces UNLINK-VL, a benchmark for evaluating cross-modal knowledge unlearning in Vision-Language Models (VLMs). The benchmark uses real-world entities from Wikidata with associated images and facts, and includes subsets for direct forgetting, relational propagation, non-target preservation, and semantic robustness. Models are trained under text-only and multimodal settings.
The paper 'Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation' proposes UNLINK-VL, a benchmark to study how unlearning in VLMs transfers across text and image modalities. It addresses the gap in cross-modal unlearning evaluation by using visually identifiable entities and their relational facts from Wikidata. The benchmark tests direct forgetting, propagation through relations, retention of related knowledge, and robustness to paraphrases, under a post-hoc setting where original training data is unavailable.
UNLINK-VL is designed to probe cross-modal consistency in unlearning by linking visual entities to one-hop and multi-hop facts. The benchmark's four subsets allow systematic evaluation of unlearning specificity and side effects, moving beyond single-modality forgetting. The post-hoc setting reflects realistic scenarios where access to original forget/retain corpora is restricted.
This work highlights the growing need for robust unlearning methods in multimodal AI systems, driven by privacy regulations and copyright concerns. As VLMs are deployed in sensitive applications, benchmarks like UNLINK-VL could become standard for auditing model compliance and safety.
For companies deploying VLMs, this benchmark offers a way to measure and demonstrate compliance with data removal requests, reducing legal and reputational risks. It could also inform the design of safer, more controllable AI products.
Next signals include adoption of UNLINK-VL in unlearning research, development of new cross-modal unlearning algorithms, and potential integration into model evaluation frameworks. Regulatory pressure may accelerate demand for verifiable unlearning in commercial VLMs.