arXiv:2604.21391cs.ROcs.AI2026-04中稿 · ICML被引 2

用意图锚定生成过程,让机器人更准更快地执行任务

From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges

论文配图:From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges
图 1 · 摘自论文原文
  • 将生成控制改为从意图出发的精细化调整
  • 在仿真和真实机器人上均实现更快收敛与更强鲁棒性
  • 适合需要精准动作控制的智能体研究者

将高层语义理解与底层物理控制相连接仍是具身智能中的核心挑战,源于认知与动作之间固有的时空尺度差异。现有生成式视觉语言动作(VLA)策略普遍采用“从噪声生成”范式,忽视了这一差异,导致表征效率低下且优化过程中条件对齐弱。本文提出ResVLA,将范式转变为“从意图精修”。认识到机器人运动天然可分解为全局意图与局部动态,ResVLA利用谱分析将控制解耦为确定性的低频锚点与随机的高频残差。通过在预测意图上锚定生成过程,模型仅聚焦于通过残差扩散桥精修局部动态。大量仿真实验表明,ResVLA在性能上具有竞争力,对语言与机器人本体扰动表现出强鲁棒性,并比标准生成基线更快收敛。该方法也在真实机器人实验中表现优异。

原文摘要 · Abstract (English)

Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action. Existing generative VLA policies typically adopt a "Generation-from-Noise" paradigm, which disregards this disparity, leading to representation inefficiency and weak condition alignment during optimization. In this work, we propose ResVLA, an architecture that shifts the paradigm to "Refinement-from-Intent." Recognizing that robotic motion naturally decomposes into global intent and local dynamics, ResVLA utilizes spectral analysis to decouple control into a deterministic low-frequency anchor and a stochastic high-frequency residual. By anchoring the generative process on the predicted intent, our model focuses strictly on refining local dynamics via a residual diffusion bridge. Extensive simulation experiments show that ResVLA achieves competitive performance, strong robustness to language and robot embodiment perturbations, and faster convergence than standard generative baselines. ResVLA also demonstrates strong performance in real-world robot experiments.

生成模型机器人控制视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。