破解视频模型性能下降之谜,用新奖励机制提升理解能力
Unhackable Temporal Rewarding for Scalable Video MLLMs
- 从强化学习视角定义时间劫持现象,提出TPL评分衡量模型对时间信息的捕捉质量
- 实验验证新框架可有效抑制模型只关注关键帧的作弊行为,显著提升视频理解效果
- 适合关注多模态大模型训练优化与视频理解任务的研究者
在追求更优视频处理多模态大模型的过程中,我们遭遇了一个令人困惑的悖论:‘反规模化定律’——数据越多、模型越大,性能反而越差。本研究揭示其根源在于‘时间劫持’:模型通过聚焦少数关键帧来投机取巧,忽视完整视频叙事。本文系统建立时间劫持的理论体系,从强化学习角度定义该现象,引入时间困惑度(Temporal Perplexity, TPL)指标评估模型对时间结构的对齐程度,并提出无懈可击的时间奖励机制(Unhackable Temporal Rewarding, UTR)以缓解此问题。理论与实证均表明,TPL是时间建模质量的可靠指标,与帧激活模式高度相关。大量实验显示,UTR不仅能有效遏制时间劫持,还能显著提升视频理解能力。本工作不仅推动了视频-人工智能系统的演进,也凸显了在多模态大模型发展中,代理奖励与真实目标对齐的重要性。
原文摘要 · Abstract (English)
In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the "anti-scaling law", where more data and larger models lead to worse performance. This study unmasks the culprit: "temporal hacking", a phenomenon where models shortcut by fixating on select frames, missing the full video narrative. In this work, we systematically establish a comprehensive theory of temporal hacking, defining it from a reinforcement learning perspective, introducing the Temporal Perplexity (TPL) score to assess this misalignment, and proposing the Unhackable Temporal Rewarding (UTR) framework to mitigate the temporal hacking. Both theoretically and empirically, TPL proves to be a reliable indicator of temporal modeling quality, correlating strongly with frame activation patterns. Extensive experiments reveal that UTR not only counters temporal hacking but significantly elevates video comprehension capabilities. This work not only advances video-AI systems but also illuminates the critical importance of aligning proxy rewards with true objectives in MLLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。