The paper proposes the Evidence-Backed Video Question Answering (E-VQA) task, requiring models to output semantic answers and precise spatiotemporal evidence (time segments and dense tracked object segmentation masks). It introduces the ST-Evidence benchmark and constructs a 160k-scale ST-Evidence-Instruct dataset. Fine-tuned Video LLMs outperform the UniPixel baseline in pixel-level localization.
The paper proposes the E-VQA task, requiring models to output answers and spatiotemporal evidence; constructs the ST-Evidence benchmark and a 160k dataset, with fine-tuned models outperforming the baseline in localization.
E-VQA extends video question answering from pure text answers to pixel-level spatiotemporal evidence, revealing a decoupling between QA accuracy and visual perception.
This work pushes video understanding from black-box reasoning toward verifiable and interpretable directions, potentially impacting fields such as video surveillance and autonomous driving that require reliable evidence.
Enhancing the trustworthiness and auditability of video AI has potential commercial value for industries requiring compliance and explainability, such as healthcare and security.
Future attention can be paid to the task's performance on more video understanding benchmarks and whether it is integrated into commercial video analysis products.