arXiv:2602.03430cs.RO2026-02

构建首个结构感知的主动响应基准,提升智能体自主决策能力

ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response

  • 基于任务图结构设计多模态主动响应框架
  • 在75项任务上实现触发检测准确率提升6.21%
  • 适合研究智能体自主性与复杂决策系统的研究者

被动智能体仅执行指令,而主动智能体能持续监测环境,以辅助和安全等高层目标为导向自主行动。但主动智能体的发展受限于缺乏专用资源。为此,我们提出ProAct-75基准,涵盖75个跨领域任务(包括协助、维护和安全监控),包含91,581条步骤级标注,并引入显式任务图,编码步骤依赖关系与并行执行可能性,为复杂决策提供结构化基础。在此基础上,我们构建ProAct-Helper基线模型,基于多模态大语言模型(MLLM)实现状态感知决策,利用任务图进行熵驱动启发式搜索,支持独立并行执行多个任务线程,而非简单模仿人类下一步。大量实验表明,ProAct-Helper优于强闭源模型:触发检测mF1提升6.21%,在线单步决策节省0.25步,并行动作率提高15.58%。

原文摘要 · Abstract (English)

While passive agents merely follow instructions, proactive agents align with higher-level objectives, such as assistance and safety by continuously monitoring the environment to determine when and how to act. However, developing proactive agents is hindered by the lack of specialized resources. To address this, we introduce ProAct-75, a benchmark designed to train and evaluate proactive agents across diverse domains, including assistance, maintenance, and safety monitoring. Spanning 75 tasks, our dataset features 91,581 step-level annotations enriched with explicit task graphs. These graphs encode step dependencies and parallel execution possibilities, providing the structural grounding necessary for complex decision-making. Building on this benchmark, we propose ProAct-Helper, a reference baseline powered by a Multimodal Large Language Model (MLLM) that grounds decision-making in state detection, and leveraging task graphs to enable entropy-driven heuristic search for action selection, allowing agents to execute parallel threads independently rather than mirroring the human's next step. Extensive experiments demonstrate that ProAct-Helper outperforms strong closed-source models, improving trigger detection mF1 by 6.21%, saving 0.25 more steps in online one-step decision, and increasing the rate of parallel actions by 15.58%.

主动智能体多模态任务图决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。