arXiv:2503.03803cs.CV2025-03CVPR被引 92

打造可记录日常的智能眼镜助手,用多模态数据提升生活效率。

EgoLife: Towards Egocentric Life Assistant

  • 用智能眼镜采集300小时第一视角生活视频,构建多视角多模态数据集。
  • 提出EgoLifeQA任务,支持长时序问答与个性化生活建议。
  • 开发EgoButler系统,实现身份识别与跨模态长上下文理解。

我们提出EgoLife项目,旨在开发一种通过AI赋能的可穿戴眼镜辅助个人提升效率的首人称生活助手。为奠定基础,我们开展全面的数据采集研究:六名参与者共同生活一周,持续使用智能眼镜进行多模态首人称视频记录,并同步获取第三人称视角参考视频,涵盖讨论、购物、烹饪、社交与娱乐等日常活动。该工作产出EgoLife数据集——一个包含300小时首人称、人际互动、多视角、多模态的日常生活数据集,附有密集标注。基于此,我们引入EgoLifeQA,一套面向生活的长上下文问答任务,用于回答如回忆过往事件、监测健康习惯及提供个性化推荐等实际问题。为应对(1)首人称视觉音频模型鲁棒性、(2)身份识别、(3)长时序信息问答等关键挑战,我们提出EgoButler系统,由EgoGPT与EgoRAG组成。EgoGPT是基于首人称数据训练的全模态模型,在首人称视频理解上达到当前最优性能;EgoRAG是检索增强组件,支持超长上下文问题回答。实验验证了其工作机制并揭示关键因素与瓶颈,指导未来改进。通过公开数据集、模型与基准,我们旨在推动首人称AI助手领域的进一步研究。

原文摘要 · Abstract (English)

We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data collection study where six participants lived together for one week, continuously recording their daily activities - including discussions, shopping, cooking, socializing, and entertainment - using AI glasses for multimodal egocentric video capture, along with synchronized third-person-view video references. This effort resulted in the EgoLife Dataset, a comprehensive 300-hour egocentric, interpersonal, multiview, and multimodal daily life dataset with intensive annotation. Leveraging this dataset, we introduce EgoLifeQA, a suite of long-context, life-oriented question-answering tasks designed to provide meaningful assistance in daily life by addressing practical questions such as recalling past relevant events, monitoring health habits, and offering personalized recommendations. To address the key technical challenges of (1) developing robust visual-audio models for egocentric data, (2) enabling identity recognition, and (3) facilitating long-context question answering over extensive temporal information, we introduce EgoButler, an integrated system comprising EgoGPT and EgoRAG. EgoGPT is an omni-modal model trained on egocentric datasets, achieving state-of-the-art performance on egocentric video understanding. EgoRAG is a retrieval-based component that supports answering ultra-long-context questions. Our experimental studies verify their working mechanisms and reveal critical factors and bottlenecks, guiding future improvements. By releasing our datasets, models, and benchmarks, we aim to stimulate further research in egocentric AI assistants.

生活助手首人称视频多模态长上下文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。