arXiv:2509.21113cs.CV2025-09被引 10

让AI视频推理过程更连贯,提升答案可信度。

MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning

  • 用动态时间规整算法设计奖励机制,约束推理过程与视频时序对齐。
  • 在自建数据集上达到87.2%准确率,通用视频基准表现也更好。
  • 适用于多种模型架构,适合关注推理一致性与可解释性的研究者。

视频推理已成为多模态大模型的关键能力,要求模型从静态感知走向对复杂场景中时序动态的连贯理解。然而现有模型常出现过程不一致问题:即使最终答案正确,中间推理仍可能偏离视频真实动态,影响可解释性与鲁棒性。为此,我们提出MOSS-ChatV,一种基于动态时间规整(DTW)的强化学习框架,通过规则化奖励使推理轨迹与时空参考对齐,实现高效过程监督而无需额外奖励模型。我们进一步将动态状态预测作为关键评估指标,构建了含标注推理轨迹的MOSS-Video基准,其中训练集用于微调MOSS-ChatV,测试集用于评估。MOSS-ChatV在MOSS-Video(测试集)上取得87.2%的准确率,并在MVBench和MMVU等通用视频基准上表现提升。该框架在Qwen2.5-VL和Phi-2等多种架构上均带来稳定增益,验证其广泛适用性。GPT-4o-as-judge的评估表明,该方法生成的推理轨迹更具一致性与稳定性。

原文摘要 · Abstract (English)

Video reasoning has emerged as a critical capability for multimodal large language models (MLLMs), requiring models to move beyond static perception toward coherent understanding of temporal dynamics in complex scenes. Yet existing MLLMs often exhibit process inconsistency, where intermediate reasoning drifts from video dynamics even when the final answer is correct, undermining interpretability and robustness. To address this issue, we introduce MOSS-ChatV, a reinforcement learning framework with a Dynamic Time Warping (DTW)-based process reward. This rule-based reward aligns reasoning traces with temporally grounded references, enabling efficient process supervision without auxiliary reward models. We further identify dynamic state prediction as a key measure of video reasoning and construct MOSS-Video, a benchmark with annotated reasoning traces, where the training split is used to fine-tune MOSS-ChatV and the held-out split is reserved for evaluation. MOSS-ChatV achieves 87.2\% on MOSS-Video (test) and improves performance on general video benchmarks such as MVBench and MMVU. The framework consistently yields gains across different architectures, including Qwen2.5-VL and Phi-2, confirming its broad applicability. Evaluations with GPT-4o-as-judge further show that MOSS-ChatV produces more consistent and stable reasoning traces.

视频推理强化学习时序对齐可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。