arXiv:2501.08771cs.CV2025-01被引 1

让模型承认不懂,能提升视频问答准确率。

Admitting Ignorance Helps the Video Question Answering Models to Answer

  • 干预问题让模型被迫承认无知,避免靠表面关联猜答案。
  • 在多选和开放题上均有效,性能显著提升。
  • 无需大改结构,适合现有模型快速升级。

视频问答(VideoQA)领域因深度学习和大规模预训练取得了显著进展。尽管模型结构复杂、视频-文本基础模型强大,但现有方法仍仅关注答案与视频-问题对的相关性最大化。我们指出,这些模型常建立捷径,导致问题与答案间存在虚假相关性,尤其当视频与文本对齐不佳时。为此,我们提出一种新训练框架:当模型面对被干预的问题时,必须承认自身无知,而非仅依赖表面的问答关联进行猜测。我们引入问题干预方法,如位移和扰动,并设计了在多选和开放题设置下让模型承认知识缺失的框架。我们将最先进的模型集成到该框架中验证效果,结果表明,该框架可在极少结构改动下显著提升视频问答模型性能。

原文摘要 · Abstract (English)

Significant progress has been made in the field of video question answering (VideoQA) thanks to deep learning and large-scale pretraining. Despite the presence of sophisticated model structures and powerful video-text foundation models, most existing methods focus solely on maximizing the correlation between answers and video-question pairs during training. We argue that these models often establish shortcuts, resulting in spurious correlations between questions and answers, especially when the alignment between video and text data is suboptimal. To address these spurious correlations, we propose a novel training framework in which the model is compelled to acknowledge its ignorance when presented with an intervened question, rather than making guesses solely based on superficial question-answer correlations. We introduce methodologies for intervening in questions, utilizing techniques such as displacement and perturbation, and design frameworks for the model to admit its lack of knowledge in both multi-choice VideoQA and open-ended settings. In practice, we integrate a state-of-the-art model into our framework to validate its effectiveness. The results clearly demonstrate that our framework can significantly enhance the performance of VideoQA models with minimal structural modifications.

视频问答模型鲁棒性认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。