首个用户自注的视角内状态数据集,助力AI理解用户意图与情绪。
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning

- 基于跨设备采集60人共180小时多模态数据,含视觉、音频、眼动等信号。
- 构建评估框架,发现现有大模型难以有效推理用户主观状态。
- 适合研究可穿戴AI、人机交互与多模态情感计算的学者使用。
尽管已有大量关于第一人称视频数据集的研究,但对用户内在状态的理解仍被忽视,而这对实现无缝的人工智能助手体验至关重要。本文提出EgoIntrospect,首个在用户主导场景下采集的第一人称数据集,包含用户自我标注的与AI助手交互的意图信息。数据通过跨设备系统同步采集,涵盖视频、音频、注视、动作及生理信号。数据集包含60名受试者的180小时记录,每人平均时长3小时。基于此,我们定义了一系列聚焦用户内在状态的任务,包括情感体验、交互意图和认知记忆。进一步处理标注以构建基准,评估现代多模态大语言模型从第一人称观察中推理用户内在状态的能力。实验表明,现有模型难以有效利用多模态信号推断用户的主观状态。数据集与标注将公开,以推动第一人称视觉与可穿戴智能助手研究。项目页面:https://ego-introspect.github.io/
原文摘要 · Abstract (English)
Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remains largely overlooked. In this work, we introduce EgoIntrospect, the first egocentric dataset captured in user-driven scenarios with self-annotations that explicitly reveal users' interactive intentions with AI assistants. EgoIntrospect was collected using a cross-device setup, providing synchronized video, audio, gaze, motion, and physiological signals. It consists of 180 hours of recordings from 60 subjects, with an average recording duration of 3 hours per subject. Leveraging EgoIntrospect, we formalize a suite of tasks centered on user internal states, including affective experience, interactive intent, and cognitive memory. We further process the annotations to construct benchmarks that evaluate the ability of modern multimodal large language models to reason about users' internal states from egocentric observations. Experiments on our benchmark suggest that existing multimodal large language models struggle to effectively leverage multimodal signals to infer users' subjective internal states. The dataset and annotations will be made publicly available to advance research in egocentric vision and wearable AI assistants. Project page: https://ego-introspect.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。