arXiv:2512.04381cs.RO2025-12被引 5

用视觉语言模型协调分离的移动与操作,提升机器人任务泛化能力

FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination

  • 将移动和操作拆分为独立策略,各自使用专属观测
  • 通过视觉语言模型共享上下文,实现跨模块协同
  • 无需人工标注阶段,自动识别任务进展,适合复杂场景

我们提出FALCON框架,用于具身智能体的定位-操作任务。该框架将移动与操作解耦为两个专用的扩散策略,分别依赖自身观测,避免单一策略融合异构信息导致的性能下降。关键创新在于利用视觉-语言基础模型将全局观察和语言指令编码为共享潜在表示,同时条件化两个扩散策略。在此基础上,引入阶段进度头,仅凭文本描述即可推断任务阶段与连续进展,无需人工标注。此外,设计协调感知对比损失,显式建模机械臂与基座动作间的兼容性。在两个需导航、精准末端执行器定位及紧密基臂协同的挑战性任务上验证,FALCON优于集中式与去中心化基线,展现出更强鲁棒性与分布外泛化能力。

原文摘要 · Abstract (English)

We present FoundAtion-model-guided decoupled LoCO-maNipulation visuomotor policies (FALCON), a framework for loco-manipulation that combines modular diffusion policies with a vision-language foundation model as the coordinator. Our approach explicitly decouples locomotion and manipulation into two specialized visuomotor policies, allowing each subsystem to rely on its own observations. This mitigates the performance degradation that arise when a single policy is forced to fuse heterogeneous, potentially mismatched observations from locomotion and manipulation. Our key innovation lies in restoring coordination between these two independent policies through a vision-language foundation model, which encodes global observations and language instructions into a shared latent embedding conditioning both diffusion policies. On top of this backbone, we introduce a phase-progress head that uses textual descriptions of task stages to infer discrete phase and continuous progress estimates without manual phase labels. To further structure the latent space, we incorporate a coordination-aware contrastive loss that explicitly encodes cross-subsystem compatibility between arm and base actions. We evaluate FALCON on two challenging loco-manipulation tasks requiring navigation, precise end-effector placement, and tight base-arm coordination. Results show that it surpasses centralized and decentralized baselines while exhibiting improved robustness and generalization to out-of-distribution scenarios.

具身智能视觉语言模型任务分解机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。