提升视频大模型对运动与时间的细粒度理解能力。
COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models

- 通过显式时间表示和外观-运动融合增强时序建模。
- 在动作和时间推理任务上平均提升4.9%和2.1%。
- 适合需要精准时序理解的视频分析场景。
视频多模态大语言模型虽有显著进展,但细粒度运动-时间理解仍不稳固。核心瓶颈不仅在于稀疏帧采样,更缺乏完整的时序建模流程,难以显式表征帧间变化、实现外观-运动交互及优化时序方向敏感性。本文提出COMET,一个基于显式时序表示、外观-运动融合与方向感知优化的时序增强框架。其架构引入基于泰勒帧差的时序运动分支,并通过时序注意力偏置增强的跨注意力将运动证据注入外观流。优化方面,结合时序先验蒸馏与前后向TC-GRPO阶段,将时序顺序直接作为学习信号,强化模型对时序运动模式的利用。实验显示,在Qwen3-VL-8B上,动作导向任务(STAR, SSv2)平均提升4.9%,时间推理任务(NExT-QA, CLEVRER, LLaVA-178K)相比BL-GRPO提升2.1%,静态感知任务(PerceptionTest)保持相当水平。该增益模式亦在InternVL2.5-8B上复现,表明COMET具备跨模型族泛化能力。
原文摘要 · Abstract (English)
Video multimodal large language models have advanced significantly, yet fine-grained motion-temporal understanding remains fragile. The core bottleneck is not only sparse frame sampling, but also the lack of a complete temporal modeling pipeline for explicitly representing frame-to-frame change, enabling appearance-motion interaction, and optimizing temporal direction sensitivity. We propose COMET, a temporally grounded framework that systematically strengthens video MLLMs through explicit temporal representation, appearance-motion fusion, and direction-aware optimization. Architecturally, COMET introduces a temporal motion branch built on Taylor frame differences and injects its motion evidence into the appearance stream via temporal attention bias-enhanced cross-attention. For optimization, COMET combines temporal prior distillation with a forward-reverse TC-GRPO stage that turns temporal order into a direct learning signal and strengthens the model's use of directional motion patterns encoded by the temporal motion branch. The method achieves consistent overall improvements with a pronounced motion-temporal bias: on Qwen3-VL-8B, action-centric tasks (STAR, SSv2) improve by 4.9% on average, temporal reasoning tasks (NExT-QA, CLEVRER, LLaVA-178K) by 2.1% over BL-GRPO, while static perception tasks (PerceptionTest) remain on par. The same gain pattern also transfers to InternVL2.5-8B, indicating that COMET generalizes across model families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。