让智能助手主动识别用户需求,而非被动等待指令。
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

- 基于多模态记忆构建上下文感知的主动决策机制。
- 在3000+服务实例上验证,支持从即时安全提醒到长期习惯指导。
- 无需训练即可适应新场景,适合可穿戴设备实时辅助。
何时智能助手应主动发言而不需用户提问?连续第一人称视频提供了丰富且动态演化的上下文,使主动式辅助成为可能。现有方法或被动等待用户询问,或对每个检测事件都响应,未考虑用户历史、当前活动及是否真正需要帮助。我们重新将主动辅助视为一种依赖上下文的决策问题:代理不仅要感知事件,还需基于累积的时间上下文判断是否干预。为此,我们提出Vinci2系统,将原有反应式助手Vinci升级为具有主动能力的本地化助手。评估方面,我们构建了首个大规模基准EgoServe,包含超过3000个服务实例,涵盖10类服务,时间跨度从即时安全警报到长期习惯指导。建模方面,我们提出EgoMemo,一种无需训练的记忆增强型代理,维护三种互补记忆表示:多尺度时间摘要、语义知识图谱和视觉嵌入存档。每个时间步,EgoMemo通过检索增强推理判断是否需要协助,并生成情境相关回应。实验表明,EgoMemo在EgoServe上建立强基线,同时在现有第一人称基准上保持竞争力。基准与代码已公开于https://sitonggong.github.io/EgoServe-page/。
原文摘要 · Abstract (English)
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user's history, current activity, or whether assistance would actually be welcome. We reframe proactive assistance as a context-dependent decision problem: the agent must not only perceive what is happening, but reason over accumulated temporal context to determine when and whether to intervene. To this end, we present Vinci2, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity. On the evaluation side, we present EgoServe, the first large-scale benchmark for proactive assistance in continuous egocentric video. EgoServe comprises over 3,000 service instances organized along 4 temporal memory horizons, ranging from immediate safety alerts to long-term habit coaching, across 10 service categories. On the modeling side, we propose EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives. At each timestep, EgoMemo performs retrieval-augmented reasoning to determine whether assistance is warranted and, if so, produces contextually grounded responses. Experiments demonstrate that EgoMemo establishes strong baselines on EgoServe while remaining competitive on existing egocentric benchmarks. Our benchmark and code are publicly available at \href{https://sitonggong.github.io/EgoServe-page/}{Vinci2}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。