通过视频推理树提升常识性视频问答的可信度。
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
- 构建视频片段的蕴含树,显式关联视觉与语言信息。
- 在多个基准上验证,显著降低模型对虚假关联的依赖。
- 适用于多种视觉语言模型,适合关注可解释性的研究者。
本文提出首个面向常识性视频问答(VQA)的视频接地蕴含树推理方法。尽管大型视觉-语言模型(VLMs)取得显著进展,但其黑箱特性及现有评测偏差导致模型易学习视频与答案间的虚假相关性。本方法通过四个步骤将VQA任务显式锚定至视频片段:蕴含树构建、视频-语言蕴含验证、树推理与动态树扩展。该方法具备跨模型、跨任务的通用性,可适配当前主流的视频与图像类VLM。为实现公平评估,我们基于大语言模型设计去偏流程,重构基准数据集的答案集以强制模型进行真实推理。在原有与去偏后基准上的系统性实验表明,本方法在不同基准、不同VLM和多种推理类型下均有效提升性能。
原文摘要 · Abstract (English)
This paper proposes the first video-grounded entailment tree reasoning method for commonsense video question answering (VQA). Despite the remarkable progress of large visual-language models (VLMs), there are growing concerns that they learn spurious correlations between videos and likely answers, reinforced by their black-box nature and remaining benchmarking biases. Our method explicitly grounds VQA tasks to video fragments in four steps: entailment tree construction, video-language entailment verification, tree reasoning, and dynamic tree expansion. A vital benefit of the method is its generalizability to current video and image-based VLMs across reasoning types. To support fair evaluation, we devise a de-biasing procedure based on large-language models that rewrites VQA benchmark answer sets to enforce model reasoning. Systematic experiments on existing and de-biased benchmarks highlight the impact of our method components across benchmarks, VLMs, and reasoning types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。