用视频生成+迭代纠错,让机器人能自适应完成复杂操作。
PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
- 通过生成候选动作视频并闭环执行,实现动态重规划。
- 首次尝试成功率仅20%-30%,但迭代修正后整体成功率超80%。
- 适用于多种机器人和视角,适合追求鲁棒性的通用机器人研发者。
我们提出PhysicalAgent,一个面向机器人操作的智能体框架,融合迭代推理、基于扩散的视频生成与闭环执行。给定文本指令后,该方法生成候选动作视频演示,在机器人上执行并针对失败进行迭代重规划,从而实现对执行错误的鲁棒恢复。我们在多种感知模态(第一人称、第三人称、仿真)和机器人形态(双臂UR3、Unitree G1人形、仿真GR1)下评估,对比最先进任务特定基线。实验表明,该方法在人类熟悉任务上最高达83%成功率。真实机器人测试显示,首次尝试成功率仅为20%-30%,但通过迭代修正,跨平台整体成功率提升至80%。结果表明,基于视频的生成式推理在通用机器人操作中具有潜力,且迭代执行对克服初始失败至关重要。本框架为可扩展、可适应、鲁棒的机器人控制开辟了新路径。
原文摘要 · Abstract (English)
We introduce PhysicalAgent, an agentic framework for robotic manipulation that integrates iterative reasoning, diffusion-based video generation, and closed-loop execution. Given a textual instruction, our method generates short video demonstrations of candidate trajectories, executes them on the robot, and iteratively re-plans in response to failures. This approach enables robust recovery from execution errors. We evaluate PhysicalAgent across multiple perceptual modalities (egocentric, third-person, and simulated) and robotic embodiments (bimanual UR3, Unitree G1 humanoid, simulated GR1), comparing against state-of-the-art task-specific baselines. Experiments demonstrate that our method consistently outperforms prior approaches, achieving up to 83% success on human-familiar tasks. Physical trials reveal that first-attempt success is limited (20-30%), yet iterative correction increases overall success to 80% across platforms. These results highlight the potential of video-based generative reasoning for general-purpose robotic manipulation and underscore the importance of iterative execution for recovering from initial failures. Our framework paves the way for scalable, adaptable, and robust robot control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。