arXiv:2605.04227cs.AIcs.HC2026-05

基于多模态感知的连续主动助手,提升长期流程任务的指导精准度。

Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks

论文配图:Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks
图 1 · 摘自论文原文
  • 通过AR眼镜持续感知动作与上下文,实时推断任务进展
  • 在真实场景中提升21%以上任务理解准确率,主动提醒精度提高2.29倍
  • 适合需要长时间流程指导的智能助手机器人场景

日常生活中普遍存在包含多个有序步骤的流程性任务。近年来,多模态大语言模型(MLLMs)已实现支持日常活动的个人助手。然而,现有系统主要提供用户提问后才响应的被动引导,或仅针对孤立短期事件的有限主动协助,难以应对长期流程任务。本文提出Pro$^2$Assist,一种能持续追踪细粒度任务进展、结合用户状态变化进行推理的步骤感知主动助手。该系统利用增强现实(AR)眼镜获取多模态数据,实现基于运动的感知;从多尺度时序动态与特定任务专家知识中提取步骤导向的流程上下文;基于感官输入与流程上下文,持续推理用户需求,并在AR眼镜上及时显示辅助信息。我们在公开来源数据集及自建测试平台收集的真实世界数据集上评估了Pro$^2$Assist。大量实验表明,其在流程动作理解准确率上优于最优基线超过21%,主动提醒时间精度最高达基线的2.29倍。20名用户的体验研究显示,90%认为该系统有用,证明其在真实场景中的有效性。

原文摘要 · Abstract (English)

Procedural tasks with multiple ordered steps are ubiquitous in daily life. Recent advances in multimodal large language models (MLLMs) have enabled personal assistants that support daily activities. However, existing systems primarily provide reactive guidance triggered by user queries, or limited proactive assistance for isolated short-term events rather than long-horizon procedural tasks. In this work, we introduce Pro$^2$Assist, a step-aware proactive assistant that continuously tracks fine-grained task progress and reasons over the user's evolving state to provide timely assistance throughout tasks. Pro$^2$Assist leverages multimodal data from augmented reality (AR) glasses to achieve motion-based perception. It then extracts step-oriented procedural context from multi-scale temporal dynamics and task-specific expert knowledge. Based on both sensory input and procedural context, Pro$^2$Assist performs continuous reasoning to infer user needs and display timely assistance on AR glasses. We evaluate Pro$^2$Assist using a dataset curated from public sources and a real-world dataset collected on our testbed with AR glasses. Extensive evaluations show that Pro$^2$Assist outperforms the best-performing baselines by over 21% in procedural action understanding accuracy, and it achieves up to 2.29x the proactive timing accuracy of baselines. A user study with 20 participants further shows that 90% find Pro$^2$Assist useful, indicating its effectiveness for real-world procedural assistance.

主动助手流程任务多模态感知AR应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。