arXiv:2509.02055cs.ROcs.AI2025-09被引 19

用统一潜空间让视觉语言动作模型更易适配新机器人和任务

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance

  • 先对齐动作空间,构建统一潜变量表示
  • 实测在真实场景中提升32%任务成功率
  • 无需大量数据,可直接用于新机器人部署

视觉-语言-动作(VLA)模型在大规模多样化数据上预训练后,在通用机器人操作中展现出巨大潜力。然而,当下游任务的机器人本体或任务内容与预训练数据不同时,动作分布差异显著,导致微调需大量数据与计算资源。为此,我们提出新型轻量级适配框架 exttt{ATE}:首先通过约束反KL散度的变分自编码器,将适配动作嵌入预训练动作潜空间的模式中,实现动作空间对齐;随后在微调过程中,利用引导机制推动扩散或流模型生成分布向目标域偏移。我们在仿真与真实世界中开展跨本体与跨任务操作实验。相比直接微调代表性VLA模型, exttt{ATE} 在仿真中平均多任务成功率提升最高达9.8%,在真实世界跨本体设置下实现32%的成功率提升。本工作为将VLA模型高效部署到新机器人平台与任务提供了通用、轻量的解决方案。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models pre-trained on large, diverse datasets show remarkable potential for general-purpose robotic manipulation. However, a primary bottleneck remains in adapting these models to downstream tasks, especially when the robot's embodiment or the task itself differs from the pre-training data. This discrepancy leads to a significant mismatch in action distributions, demanding extensive data and compute for effective fine-tuning. To address this challenge, we introduce \textbf{Align-Then-stEer (\texttt{ATE})}, a novel, data-efficient, and plug-and-play adaptation framework. \texttt{ATE} first aligns disparate action spaces by constructing a unified latent space, where a variational autoencoder constrained by reverse KL divergence embeds adaptation actions into modes of the pre-training action latent distribution. Subsequently, it steers the diffusion- or flow-based VLA's generation process during fine-tuning via a guidance mechanism that pushes the model's output distribution towards the target domain. We conduct extensive experiments on cross-embodiment and cross-task manipulation in both simulation and real world. Compared to direct fine-tuning of representative VLAs, our method improves the average multi-task success rate by up to \textbf{9.8\%} in simulation and achieves a striking \textbf{32\% success rate gain} in a real-world cross-embodiment setting. Our work presents a general and lightweight solution that greatly enhances the practicality of deploying VLA models to new robotic platforms and tasks.

机器人操作视觉语言动作模型适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。