打造可主动纠错的流程助手,让系统实时引导并恢复偏离步骤的操作。
Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance

- 分离规划与交互的架构,支持实时判断何时干预及如何指导。
- 在6个数据集上超越闭源和开源模型,尤其在偏离流程时恢复能力更强。
- 新基准涵盖真实错乱场景,适合做流程指导与自适应交互研究。
我们提出一个主动式多模态助手系统,能在用户执行流程任务时实时提供分步指导,并自主决定何时打断及如何辅导。当前进展受限于缺乏大规模、跨领域的现实场景基准,尤其是用户偏离预期步骤的情况。为此,我们贡献四方面:(1) 发布大尺度可穿戴视角数据集 EgoProactive,包含显式的离计划(OOP)标注与恢复步骤;(2) 将五个现有基准(Ego4D、EPIC-KITCHENS、EgoExo4D、HoloAssist、HowTo100M)统一整合为 Pro²Bench,采用一致的主动引导范式;(3) 提出解耦式规划-交互架构,专精于流程状态、视觉线索与恢复注入;(4) 引入跨模型家族的后训练方法,在 Llama-4 与 Qwen-3.6-VL 上验证跨骨干复制能力。大量实验表明,训练后的 Llama-4 系统在六大数据集上显著优于强闭源基线(Claude Opus 4.6、Gemini 3.1 Pro、GPT 5.2)和开源基线(Qwen3 VL 235B)的客观干预质量。基于最优计划的实验进一步显示,当计划质量可控时,该双路模型能生成高质量引导,并在处理离计划情况时取得显著提升。
原文摘要 · Abstract (English)
We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \textit{how} to coach. However, progress is limited by the absence of large-scale, cross-domain benchmarks that reflect realistic conditions, particularly the common case in which users deviate from the expected step sequence. We address this gap with four contributions: \textbf{(1)}~we release \textbf{EgoProactive}, a large-scale wearable-egocentric dataset for proactive procedural assistance with explicit Out-of-Plan (OOP) annotations and recovery steps; \textbf{(2)}~we augment five established benchmarks (Ego4D, EPIC-KITCHENS, EgoExo4D, HoloAssist, HowTo100M) into \textbf{Pro\textsuperscript{2}Bench} under a unified proactive-guidance schema; \textbf{(3)}~we propose a \textbf{decoupled planner--interaction architecture} specialized for procedural state, visual cues, and recovery injection; \textbf{(4)}~we introduce a post-training recipe that transfers across model families, validated by cross-backbone replication on Llama~4 and Qwen-3.6-VL. In extensive experiments, our trained Llama-4 system substantially improves objective intervention quality over strong proprietary baselines (Claude Opus~4.6, Gemini~3.1~Pro, GPT~5.2) and open-weight baselines (Qwen3~VL~235B) baselines across all six datasets. Oracle-plan experiments further show that, when plan quality is controlled, the trained duplex model produces high-quality guidance and large gains on Out-of-Plan recovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。