让机器人通过语言指令生成导航视频并实时控制运动,成功率超60%。
Action Agent: Agentic Video Generation Meets Flow-Constrained Diffusion

- 用大模型动态选模型、优化提示词,提升视频生成成功率至86%
- 用流约束扩散模型将视频转为连续速度指令,实现在真实机器人上64.7%任务完成率
- 适合作为多机器人形态通用的语义导航系统参考
我们提出Action Agent,一种两阶段框架,将代理式导航视频生成与流约束扩散控制结合,实现多具身机器人导航。第一阶段中,大语言模型作为调度模块,选择视频扩散模型,通过迭代验证优化提示词,并积累跨任务记忆,从语言和图像输入合成物理合理的第一人称导航视频,使视频生成成功率从35%(单次)提升至86%(50个导航任务)。第二阶段引入FlowDiT——一种流约束扩散变换器,将优化后的目标视频与语言指令转化为连续速度指令,采用DINOv2视觉特征、学习到的光流表示自我运动,以及CLIP语言嵌入实现语义停止。在RECON室外导航数据集上预训练,在Isaac Sim中203个Unitree G1人形机器人片段上微调以校准速度动力学。仅用4300万参数的单一检查点,在模拟中实现73.2%导航成功率,真实Unitree G1在未见过的室内环境中开环执行下达到64.7%任务完成率,运行频率为40–47 Hz。我们在三种具身形态上评估:Unitree G1人形机器人(真实硬件)、无人机和轮式移动机器人(Isaac Sim),证明解耦轨迹想象与执行可构建可扩展、具身感知的语言引导导航范式。
原文摘要 · Abstract (English)
We present Action Agent, a two-stage framework that unifies agentic navigation video generation with flow-constrained diffusion control for multi-embodiment robot navigation. In Stage I, a large language model (LLM) acts as an orchestration module that selects video diffusion models, refines prompts through iterative validation, and accumulates cross-task memory to synthesize physically plausible first-person navigation videos from language and image inputs. This increases video generation success from 35% (single-shot) to 86% across 50 navigation tasks. In Stage II, we introduce FlowDiT, a Flow-Constrained Diffusion Transformer that converts optimized goal videos and language instructions into continuous velocity commands using action-space denoising diffusion. FlowDiT integrates DINOv2 visual features, learned optical flow for ego-motion representation, and CLIP language embeddings for semantic stopping. We pretrain on the RECON outdoor navigation dataset and fine-tune on 203 Unitree G1 humanoid episodes collected in Isaac Sim to calibrate velocity dynamics. A single 43M-parameter checkpoint achieves 73.2% navigation success in simulation and 64.7% task completion on a real Unitree G1 in unseen indoor environments under open-loop execution, while operating at 40--47 Hz. We evaluate Action Agent across three embodiments: a Unitree G1 humanoid (real hardware), a drone, and a wheeled mobile robot (Isaac Sim), demonstrating that decoupling trajectory imagination from execution yields a scalable and embodiment-aware paradigm for language-guided navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。