arXiv:2506.09079cs.CVcs.AI2025-06被引 6

用中间任务让视频模型同时精通问答与描述

VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks

  • 设计两个中间代理任务,引导模型兼顾发散理解与收敛推理
  • 在QA和描述任务上均显著提升,实现单一模型双优表现
  • 适合需要通用视频理解能力的研究者与开发者

“先推理再回答”范式结合强化学习,在多模态大模型中展现出巨大潜力。然而其在视频领域的应用导致模型专精于问答(QA)或描述(captioning)任务之一,难以兼顾二者。直接合并两类任务的奖励信号会引发性能相互损害,我们归因于两者任务本质的冲突。为此,提出一种新训练框架,包含两个中间代理任务:DarkEventInfer通过遮蔽视频片段,要求模型基于上下文推断被遮内容;MixVidQA则将两段不同视频交错呈现,挑战模型分离并聚焦其中一段进行推理。这两个任务促使模型同时发展全局发散理解与精准收敛推理能力。基于此框架,提出VidBridge-R1,首个能有效弥合范式冲突的通用视频推理模型。大量实验表明,VidBridge-R1在单个模型中显著提升QA与描述性能,验证了该方法在构建更通用、更强视频理解模型方面的有效性。

原文摘要 · Abstract (English)

The "Reason-Then-Respond" paradigm, enhanced by Reinforcement Learning, has shown great promise in advancing Multimodal Large Language Models. However, its application to the video domain has led to specialized models that excel at either question answering (QA) or captioning tasks, but struggle to master both. Naively combining reward signals from these tasks results in mutual performance degradation, which we attribute to a conflict between their opposing task natures. To address this challenge, we propose a novel training framework built upon two intermediate proxy tasks: DarkEventInfer, which presents videos with masked event segments, requiring models to infer the obscured content based on contextual video cues; and MixVidQA, which presents interleaved video sequences composed of two distinct clips, challenging models to isolate and reason about one while disregarding the other. These proxy tasks compel the model to simultaneously develop both holistic, divergent understanding and precise, convergent reasoning capabilities. Embodying this framework, we present VidBridge-R1, the first versatile video reasoning model that effectively bridges the paradigm conflict. Extensive experiments show that VidBridge-R1 achieves significant performance gains on both QA and captioning within one model, demonstrating the efficacy of our approach in fostering more generalizable and powerful video understanding models.

视频理解强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。