arXiv:2503.20472cs.CVcs.AI2025-03ICCV被引 10

通过多视角采样与自奖励机制提升长视频理解准确率

From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment

  • 采用分箱采样策略生成多组关键帧组合,丰富视觉上下文
  • 在7个数据集上显著提升3种大模型的长视频问答准确率
  • 适合需要高精度长视频分析的研究者和应用开发者

多模态大语言模型在视频理解方面表现优异,但长视频理解仍具挑战,因模型单次推理只能处理有限帧数,可能遗漏关键视觉信息。为此,我们提出通过视觉上下文采样生成多个预测,再通过评分机制选择最终答案。具体而言,设计分箱采样策略,使大模型基于不同关键帧组合生成多样化回答,从而增强视觉上下文。为从采样结果中确定最终答案,采用线性组合三项得分的自奖励机制:(1) 频率得分,反映各选项出现频率;(2) 边际置信度得分,体现模型预测的组间与组内一致性;(3) 针对不同问题类型设计的推理得分,包括全局问题的线索引导回答与局部问题的时间自聚焦。频率得分确保多数正确性,置信度得分反映预测可靠性,类型化推理得分针对关键视觉信息稀疏场景采用定制策略。实验表明,该方法在7个数据集上覆盖了大量长视频问题的正确答案,显著提升了3种多模态大模型的性能。

原文摘要 · Abstract (English)

Multi-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process a finite number of frames in a single inference, potentially omitting crucial visual information. To address the challenge, we propose generating multiple predictions through visual context sampling, followed by a scoring mechanism to select the final prediction. Specifically, we devise a bin-wise sampling strategy that enables MLLMs to generate diverse answers based on various combinations of keyframes, thereby enriching the visual context. To determine the final prediction from the sampled answers, we employ a self-reward by linearly combining three scores: (1) a frequency score indicating the prevalence of each option, (2) a marginal confidence score reflecting the inter-intra sample certainty of MLLM predictions, and (3) a reasoning score for different question types, including clue-guided answering for global questions and temporal self-refocusing for local questions. The frequency score ensures robustness through majority correctness, the confidence-aligned score reflects prediction certainty, and the typed-reasoning score addresses cases with sparse key visual information using tailored strategies. Experiments show that this approach covers the correct answer for a high percentage of long video questions, on seven datasets show that our method improves the performance of three MLLMs.

长视频理解多模态模型自奖励机制关键帧采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。