arXiv:2503.24376cs.CVcs.AI2025-03被引 25

用强化学习提升视频理解,发现模型更会看但逻辑变差

Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1

  • 用强化学习优化多模态大模型的视频理解能力
  • 在多种场景下均优于监督微调,尤其在跨环境任务表现突出
  • 适合关注视频推理与强化学习结合的研究者

链式思维(COT)生成显著提升了大语言模型(LLM)的推理能力,强化学习(RL)作为后训练方法表现出色。多模态大语言模型(MLLM)虽具备推理潜力,但在需感知与逻辑推理结合的任务中仍研究不足。为此,我们提出SEED-Bench-R1,一个系统评估MLLM后训练方法在视频理解中的基准。该基准包含复杂现实视频和日常规划类多选题,要求精细感知与推理。评估涵盖三层次泛化:分布内、跨环境、跨环境-任务,并配备大规模训练数据集与可验证真值答案。以Qwen2-VL-Instruct-7B为基线模型,对比RL与监督微调(SFT),结果显示RL在数据效率和分布内外任务上均表现更优,甚至超越SFT在LongVideoBench等通用视频理解基准上的表现。深入分析发现,RL增强视觉感知能力,但推理链逻辑连贯性下降。识别出推理不一致、忽略视觉线索等关键局限,并建议未来改进基座模型推理、奖励建模及对噪声信号的鲁棒性。

原文摘要 · Abstract (English)

Recent advancements in Chain of Thought (COT) generation have significantly improved the reasoning capabilities of Large Language Models (LLMs), with reinforcement learning (RL) emerging as an effective post-training approach. Multimodal Large Language Models (MLLMs) inherit this reasoning potential but remain underexplored in tasks requiring both perception and logical reasoning. To address this, we introduce SEED-Bench-R1, a benchmark designed to systematically evaluate post-training methods for MLLMs in video understanding. It includes intricate real-world videos and complex everyday planning tasks in the format of multiple-choice questions, requiring sophisticated perception and reasoning. SEED-Bench-R1 assesses generalization through a three-level hierarchy: in-distribution, cross-environment, and cross-environment-task scenarios, equipped with a large-scale training dataset with easily verifiable ground-truth answers. Using Qwen2-VL-Instruct-7B as a base model, we compare RL with supervised fine-tuning (SFT), demonstrating RL's data efficiency and superior performance on both in-distribution and out-of-distribution tasks, even outperforming SFT on general video understanding benchmarks like LongVideoBench. Our detailed analysis reveals that RL enhances visual perception but often produces less logically coherent reasoning chains. We identify key limitations such as inconsistent reasoning and overlooked visual cues, and suggest future improvements in base model reasoning, reward modeling, and RL robustness against noisy signals.

视频理解强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。