用大模型生成伪标签,让关键帧选择更准更高效。
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
- 用大模型生成伪关键帧标签,解决监督信号稀疏问题。
- 在NExT-QA上提升时序与因果类问题准确率,效果显著。
- 适合需要高效视频推理的场景,如智能客服、自动摘要。
大型多模态模型(LMMs)在视频问答(VideoQA)中表现优异,但因推理成本高和信息稀释,仍面临挑战。关键帧选择可提升效率并增强推理能力,但仅依赖图像-文本相似度会导致监督信号稀疏和冗余帧选取。本文提出一种问题感知的关键帧选择框架,包含两个组件:由LMM生成的伪关键帧标签提供有效监督,以及覆盖正则化机制以促进时间维度上的多样性和互补性证据。在NExT-QA数据集上的实验表明,该方法显著提升了准确率,尤其在时序和因果类问题上表现突出,验证了关键帧选择作为可学习模块在VideoQA中的有效性。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) have recently demonstrated remarkable performance in video question answering (VideoQA), yet reasoning over video remains challenging due to high inference cost and diluted information. Keyframe selection offers efficiency and sharper reasoning but suffers from sparse supervision and redundant frame choices when relying only on image-text similarity. We present a question-aware keyframe selection framework with two components: pseudo keyframe labels derived from LMMs that provide informative supervision and a coverage regularization that promotes diverse, complementary evidence across time. Experiments on NExT-QA show that our method significantly improves accuracy, especially for temporal and causal question types, establishing keyframe selection as an effective and learnable module for VideoQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。