arXiv:2510.26113cs.CVcs.AI2025-10中稿 · ECCV被引 4

测试视频大模型在不同视角下的时间理解一致性,发现现有模型表现不稳定。

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

  • 构建同步第一人称与第三人称视频对的基准数据集
  • 发现模型跨视角时间理解一致性显著低于单视角表现
  • 提出新强化学习方法提升多视角一致性,效果优于传统微调

视频大模型在不同视角下对同一事件的时间理解是否一致?为研究此问题,我们引入EgoExo-Con(一致性)基准,包含同步的第一人称与第三人称视频对及人工标注的问题,确保所有概念在两个视角中均可见。该基准聚焦两类时间理解任务:时间验证与时间定位,评估不仅关注正确性,还强调视角间的一致性。分析揭示现有视频大模型存在两大缺陷:(1) 模型难以维持一致性,跨视角表现远低于单视角;(2) 用双视角同步视频直接微调后,虽一致性提升但常劣于单视角训练模型。为此,我们提出View-GRPO,一种新型强化学习框架,有效增强视角特异性时间推理,同时促进跨视角理解一致性。实验表明,该方法在提升跨视角一致性方面表现优异。所有资源已公开于 https://minjoong507.github.io/projects/EgoExo-Con/

原文摘要 · Abstract (English)

Do Video-LLMs have consistent temporal understanding when videos capture the same event from different viewpoints? To study this question, we introduce EgoExo-Con(sistency), a benchmark of synchronized egocentric and exocentric video pairs with human-refined queries that ensure all concepts are visible in both viewpoints. EgoExo-Con emphasizes two temporal understanding tasks: Temporal Verification and Temporal Grounding. It evaluates not only correctness but consistency across viewpoints. Our analysis reveals two critical limitations of existing Video-LLMs: (1) models often fail to maintain consistency, with results far worse than their single-view performances. (2) When naively finetuned with synchronized videos of both viewpoints, the models show improved consistency but often underperform those trained on a single view. For improvements, we propose View-GRPO, a novel reinforcement learning framework that effectively strengthens view-specific temporal reasoning while encouraging consistent comprehension across viewpoints. Our method demonstrates its superior temporal understanding capabilities, especially for improving cross-view consistency. All resources have been made available at https://minjoong507.github.io/projects/EgoExo-Con/

视频理解多视角一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。