首个跨第一/第三人称视频理解基准,评测模型跨视角推理能力
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
- 构建11个子任务的跨视角视频理解基准
- 13个主流多模态大模型在跨视角对齐上表现不佳
- 适合研究具身智能与类人辅助系统的团队使用
将第一人称(主体视角)与第三人称(客体视角)之间的知识迁移与整合是人类智能的核心,使人能够从他人经验中学习并传达自身体验。尽管多模态大语言模型(MLLMs)发展迅速,其跨视角推理能力仍缺乏研究。为此,我们提出EgoExoBench,首个面向主体-客体视角视频理解与推理的基准。该基准基于公开数据集构建,包含超过7,300个问答对,涵盖十一项子任务,分为三大核心挑战:语义对齐、视角关联和时间推理。我们评估了13个最先进的MLLMs,发现这些模型虽在单视角任务上表现优异,但在跨视角语义对齐、视角准确关联及主体-客体情境下的时序推断方面存在显著短板。我们期望EgoExoBench能为追求类人跨视角智能的具身智能体与智能助手研究提供重要支持。
原文摘要 · Abstract (English)
Transferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own experiences. Despite rapid progress in multimodal large language models (MLLMs), their ability to perform such cross-view reasoning remains unexplored. To address this, we introduce EgoExoBench, the first benchmark for egocentric-exocentric video understanding and reasoning. Built from publicly available datasets, EgoExoBench comprises over 7,300 question-answer pairs spanning eleven sub-tasks organized into three core challenges: semantic alignment, viewpoint association, and temporal reasoning. We evaluate 13 state-of-the-art MLLMs and find that while these models excel on single-view tasks, they struggle to align semantics across perspectives, accurately associate views, and infer temporal dynamics in the ego-exo context. We hope EgoExoBench can serve as a valuable resource for research on embodied agents and intelligent assistants seeking human-like cross-view intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。