构建首个真实场景下持续运行的多模态助手机能评测基准
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
- 设计14小时同步视频/音频/文本数据流,覆盖四大生活场景
- 定义12项诊断任务,包含3291个经人工验证的问题
- 提出实时准确率与记忆持久时间新指标,支持长时序评估
现实世界中,第一人称智能助手需处理多模态输入(视频、音频、文本),实时响应并维持动态长时记忆。然而现有评测基准通常孤立评估各项能力,缺乏真实流式场景或仅支持短期任务。本文提出 extbf{TeleEgo}:一个长期、流式、全模态的基准,用于评估真实日常情境下的第一人称智能助手。数据集包含每位参与者超过14小时的同步第一人称视频、音频与文本,覆盖工作学习、生活方式、社交活动及外出文化四大领域。所有数据对齐统一全局时间线,并包含高质量视觉叙述与语音转录,经人工精修。TeleEgo定义了12项跨三大核心能力的诊断任务:记忆(回忆过往事件)、理解(解析当前情境)、跨记忆推理(关联遥远事件)。共含3,291个经人工验证的问答项,涵盖单选、二元、多选与开放题,严格在流式环境下评估。我们提出实时准确率(RTA)以联合衡量正确性与响应速度,以及记忆持久时间(MPT)作为面向未来的长时保留度指标。本研究报告当前模型的RTA结果,并发布TeleEgo及配套的MPT评估框架,为未来具备更强流式记忆的智能助手提供真实、可扩展的评测基准,推动对实时行为与长周期记忆的系统研究。
原文摘要 · Abstract (English)
Egocentric AI assistants in real-world settings must process multi-modal inputs (video, audio, text), respond in real time, and retain evolving long-term memory. However, existing benchmarks typically evaluate these abilities in isolation, lack realistic streaming scenarios, or support only short-term tasks. We introduce \textbf{TeleEgo}, a long-duration, streaming, omni-modal benchmark for evaluating egocentric AI assistants in realistic daily contexts. The dataset features over 14 hours per participant of synchronized egocentric video, audio, and text across four domains: work \& study, lifestyle \& routines, social activities, and outings \& culture. All data is aligned on a unified global timeline and includes high-quality visual narrations and speech transcripts, curated through human refinement.TeleEgo defines 12 diagnostic subtasks across three core capabilities: Memory (recalling past events), Understanding (interpreting the current moment), and Cross-Memory Reasoning (linking distant events). It contains 3,291 human-verified QA items spanning multiple question formats (single-choice, binary, multi-choice, and open-ended), evaluated strictly in a streaming setting. We propose Real-Time Accuracy (RTA) to jointly capture correctness and responsiveness under tight decision windows, and Memory Persistence Time (MPT) as a forward-looking metric for long-term retention in continuous streams. In this work, we report RTA results for current models and release TeleEgo, together with an MPT evaluation framework, as a realistic and extensible benchmark for future egocentric assistants with stronger streaming memory, enabling systematic study of both real-time behavior and long-horizon memory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。