用多任务强化学习提升大模型对视频时间的理解能力。
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
- 设计跨任务奖励机制,区分三种时间区间对应关系。
- 在多个基准上达到当前最优,单任务与泛化性能同步提升。
- 适合研究视频时序理解或大模型训练的学者参考。
提升多模态大语言模型(MLLMs)的时间理解能力对于长视频分析至关重要,支持时间定位、动作检测和时敏问答等任务。尽管强化学习(RL)已被用于改进时间推理,但现有方法通常局限于有限的任务类型和数据,限制了其在多样化时间理解场景中的泛化能力。为此,我们提出 TempR1,一种面向时间感知的多任务强化学习框架,系统性增强 MLLMs 的时间理解能力。我们构建了一个包含多种时间结构与语义的多任务语料库,并基于组相对策略优化(GRPO)算法实现稳定高效的跨任务优化。具体而言,我们将时间任务分为三类预测区间与真实实例的对应关系,并为每类设计定制化的定位奖励,使 TempR1 能捕捉细粒度的时间依赖性并适应不同时间模式。大量实验表明,TempR1 在多个基准上均取得领先性能。其在互补任务上的联合优化产生显著协同效应,同时提升泛化能力与单任务表现,为 MLLMs 中的时间推理提供了可扩展且原理清晰的范式。
原文摘要 · Abstract (English)
Enhancing the temporal understanding of Multimodal Large Language Models (MLLMs) is essential for advancing long-form video analysis, enabling tasks such as temporal localization, action detection, and time-sensitive question answering. While reinforcement learning (RL) has recently been explored for improving temporal reasoning, existing approaches are often confined to limited task types and data, restricting their generalization across diverse temporal understanding scenarios. To address this challenge, we present TempR1, a temporal-aware multi-task reinforcement learning framework that systematically strengthens MLLMs' temporal comprehension. We curate a multi-task corpus that exposes the model to diverse temporal structures and semantics, and build upon the Group Relative Policy Optimization (GRPO) algorithm to achieve stable and effective cross-task optimization. Specifically, we categorize temporal tasks into three correspondence types between predicted intervals and ground-truth instances, and design tailored localization rewards for each, enabling TempR1 to capture fine-grained temporal dependencies and adapt to different temporal patterns. Extensive experiments demonstrate that TempR1 attains state-of-the-art performance across multiple benchmarks. Moreover, its joint optimization over complementary tasks yields a strong synergistic effect, enhancing both generalization and single-task performance, establishing a scalable and principled paradigm for temporal reasoning in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。