arXiv:2511.12530cs.CV2025-11AAAI被引 5

用因果约束选关键帧,提升视频理解精度

ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding

  • 基于因果信息瓶颈,同时满足预测充分性和因果必要性
  • 在有限帧条件下,多个数据集上超越现有方法
  • 适合需要高效视频分析的场景,如智能监控与内容推荐

由于视觉语言模型输入令牌数量有限,且视频中相关信息具有时间稀疏性,关键帧选择对视频理解至关重要。现有方法多关注信息量,却忽略因果决定性。本文提出ReaSon框架,将关键帧选择建模为优化问题,引入新型因果信息瓶颈(CIB),明确将关键帧定义为同时满足预测充分性和因果必要性的帧。ReaSon通过可学习策略网络从视觉相关候选帧中选择关键帧以实现预测充分性,并利用反事实干预评估因果必要性。设计与CIB原则一致的复合奖励,通过强化学习优化选择策略。在NExT-QA、EgoSchema和Video-MME三个数据集上的实验表明,在有限帧设置下,ReaSon持续优于现有最先进方法,验证了其有效性与泛化能力。

原文摘要 · Abstract (English)

Keyframe selection has become essential for video understanding with vision-language models (VLMs) due to limited input tokens and the temporal sparsity of relevant information across video frames. Video understanding often relies on effective keyframes that are not only informative but also causally decisive. To this end, we propose Reinforced Causal Search with Information Bottleneck (ReaSon), a framework that formulates keyframe selection as an optimization problem with the help of a novel Causal Information Bottleneck (CIB), which explicitly defines keyframes as those satisfying both predictive sufficiency and causal necessity. Specifically, ReaSon employs a learnable policy network to select keyframes from a visually relevant pool of candidate frames to capture predictive sufficiency, and then assesses causal necessity via counterfactual interventions. Finally, a composite reward aligned with the CIB principle is designed to guide the selection policy through reinforcement learning. Extensive experiments on NExT-QA, EgoSchema, and Video-MME demonstrate that ReaSon consistently outperforms existing state-of-the-art methods under limited-frame settings, validating its effectiveness and generalization ability.

视频理解关键帧因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。