arXiv:2501.06835cs.CV2025-01EMNLP被引 16

构建超长第一视角视频理解基准,填补长时行为分析空白

X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding

  • 用真实数据+合成日程生成432段超长第一视角视频
  • 最长视频达16.4小时,现有模型表现普遍较差
  • 适合研究长期行为理解与记忆建模的学者使用

长时第一视角视频理解能提供丰富的上下文信息,为具身智能、长期行为分析和个性化辅助技术带来重要价值。然而,现有基准数据集多聚焦于单个短时(数分钟至数十分钟)或中等长度视频,缺乏对超长第一视角记录的评估能力。为此,我们提出X-LeBench,一个专为极端长时第一视角视频理解设计的新基准。该数据集通过生命日志仿真流程,生成与真实视频数据一致的连贯日常计划,并将合成计划与Ego4D——一个涵盖广泛日常生活场景的大规模第一视角视频数据集——的真实影像灵活融合,形成432段模拟视频日志,时长从23分钟到16.4小时不等。对多个基线系统和多模态大语言模型(MLLMs)的评估显示其性能普遍不佳,凸显了时间定位、推理、上下文聚合和记忆保持等核心挑战,表明亟需更先进的模型架构。

原文摘要 · Abstract (English)

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activity analysis, and personalized assistive technologies. However, existing benchmark datasets primarily focus on single, short (\eg, minutes to tens of minutes) to moderately long videos, leaving a substantial gap in evaluating extensive, ultra-long egocentric video recordings. To address this, we introduce X-LeBench, a novel benchmark dataset meticulously designed to fill this gap by focusing on tasks requiring a comprehensive understanding of extremely long egocentric video recordings. Our X-LeBench develops a life-logging simulation pipeline that produces realistic, coherent daily plans aligned with real-world video data. This approach enables the flexible integration of synthetic daily plans with real-world footage from Ego4D-a massive-scale egocentric video dataset covers a wide range of daily life scenarios-resulting in 432 simulated video life logs spanning from 23 minutes to 16.4 hours. The evaluations of several baseline systems and multimodal large language models (MLLMs) reveal their poor performance across the board, highlighting the inherent challenges of long-form egocentric video understanding, such as temporal localization and reasoning, context aggregation, and memory retention, and underscoring the need for more advanced models.

视频理解长时序列第一视角基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。