arXiv:2506.03097cs.CVcs.AI2025-06被引 12

EgoVLM通过强化学习提升第一视角视频理解能力,无需思维链标注。

EgoVLM: Policy Optimization for Egocentric Video Understanding

  • 用分组相对策略优化法直接训练,跳过思维链监督阶段。
  • 在EgoSchema上比Qwen2.5-VL 3B/7B高14.33/13.87准确率点。
  • 生成推理过程增强可解释性,适合需要透明决策的应用。

新兴的具身AI应用(如可穿戴相机、自主代理)凸显了从第一视角视频流中进行鲁棒推理的需求。我们提出EgoVLM,一种专为第一视角视频场景设计的视觉语言模型,融合视觉理解与时空推理。EgoVLM采用组相对策略优化(GRPO)进行微调,该强化学习方法可对齐模型输出与人类推理步骤。借鉴DeepSeek R1-Zero的方法,我们直接通过强化学习训练,无需任何链式思维(CoT)数据的监督微调阶段。在第一视角视频问答基准上评估,领域特定训练显著优于通用视觉语言模型。我们的EgoVLM-3B仅在非CoT第一视角数据上训练,便在EgoSchema基准上分别超过基线Qwen2.5-VL 3B和7B模型14.33和13.87准确率点。通过显式生成推理轨迹,EgoVLM提升了可解释性,适用于下游应用。此外,我们提出一种基于关键帧的奖励机制,结合显著帧选择以指导强化学习优化。该奖励设计为未来时序锚定的第一视角推理研究开辟了新方向。

原文摘要 · Abstract (English)

Emerging embodied AI applications, such as wearable cameras and autonomous agents, have underscored the need for robust reasoning from first person video streams. We introduce EgoVLM, a vision-language model specifically designed to integrate visual comprehension and spatial-temporal reasoning within egocentric video contexts. EgoVLM is fine-tuned via Group Relative Policy Optimization (GRPO), a reinforcement learning method adapted to align model outputs with human-like reasoning steps. Following DeepSeek R1-Zero's approach, we directly tune using RL without any supervised fine-tuning phase on chain-of-thought (CoT) data. We evaluate EgoVLM on egocentric video question answering benchmarks and show that domain-specific training substantially improves performance over general-purpose VLMs. Our EgoVLM-3B, trained exclusively on non-CoT egocentric data, outperforms the base Qwen2.5-VL 3B and 7B models by 14.33 and 13.87 accuracy points on the EgoSchema benchmark, respectively. By explicitly generating reasoning traces, EgoVLM enhances interpretability, making it well-suited for downstream applications. Furthermore, we introduce a novel keyframe-based reward that incorporates salient frame selection to guide reinforcement learning optimization. This reward formulation opens a promising avenue for future exploration in temporally grounded egocentric reasoning.

第一视角视觉语言模型强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。