arXiv:2508.07033cs.RO2025-08中稿 · RSS 2026 Workshop …被引 2

提出统一框架P³,让机器人更灵活地感知环境、用工具和规划多任务。

$\mathcal{P}^3$: Toward Versatile Embodied Agents

  • 主动感知环境,无需依赖工具反馈
  • 支持无反馈插拔工具,动态调整任务顺序
  • 实测验证通用性,适合真实场景部署

具身智能体在与物理环境交互方面已展现潜力,但通用型智能体仍面临三大瓶颈:动态环境感知、开放工具接入和复杂多任务规划。现有方法完全依赖工具反馈来追踪场景变化和任务进展,导致实时适应性差、错误累积严重且工具兼容性有限;多任务调度因难以处理任务依赖与优先级冲突而研究不足。为此,我们提出 $/mathcal{P}^3$,一个融合实时感知与动态调度的统一框架,能主动获取任务相关环境信息,无需反馈即可插拔并使用工具,并通过优先处理紧急任务、根据依赖关系动态调整顺序来规划多任务执行。我们还构建了主动任务感知(ATP)基准,定量评估视觉语言模型在主动场景理解与任务提议方面的能力。ATP基准测试表明多个VLM可检测并提出主动任务;真实机器人实验全面验证该方法缩小了基准与实际部署间的差距,实现了可迁移的通用具身智能体。代码与数据见 https://github.com/fz-zsl/P3。

原文摘要 · Abstract (English)

Embodied agents have demonstrated promising capabilities in interacting with physical environments. Yet, versatile embodied agents face three core bottlenecks: dynamic environmental perception, open tool access, and complex multi-task planning. Prior methods depend entirely on tool feedback to track scene changes and task progress, leading to poor real-time adaptability, error accumulation, and limited tool compatibility; multi-task scheduling is also understudied due to the difficulty of handling task dependencies and conflicting priorities. To address these limitations, we propose $\mathcal P^3$, a unified framework integrating real-time perception and dynamic scheduling, which perceives task-relevant information actively from the environment, plugs and utilizes tools without feedback requirements, and plans multi-task execution by prioritizing urgent tasks and dynamically adjusting task order based on dependencies. We additionally build the Active Task Perception (ATP) benchmark to quantitatively measure VLMs' capacity for active scene understanding and task proposal. Evaluations on the ATP benchmark verify that multiple VLMs can detect and propose active tasks, and comprehensive real-world robot experiments prove our method bridges the gap between benchmarks and practical deployment, yielding transferable general-purpose embodied agents. Code and data are available at https://github.com/fz-zsl/P3.

具身智能多任务规划视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。