arXiv:2604.10055cs.RO2026-04被引 4

提出分阶段训练框架,让视觉语言动作模型更抗干扰。

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations

论文配图:STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations
图 1 · 摘自论文原文
  • 分两阶段训练:先学抗干扰,再恢复任务精度
  • 在多个机器人基准上提升成功率,最多增16.49%
  • 适合需要稳定运行的现实机器人系统

尽管视觉-语言-动作(VLA)模型在具身任务中表现优异,但在多模态扰动下仍极脆弱,视觉退化与语言噪声共同导致分布偏移,降低任务执行效果。现有方法依赖联合训练扰动数据,将鲁棒性视为静态目标,导致鲁棒性与任务精度优化冲突。本文提出STRONG-VLA,一种解耦微调框架,显式分离鲁棒性学习与任务对齐优化。第一阶段通过渐进式多模态扰动课程训练,实现受控分布偏移下的渐进鲁棒性学习;第二阶段在干净任务分布下重新对齐,恢复执行精度并保持鲁棒性。我们构建了涵盖28种扰动类型的综合基准,覆盖真实传感器噪声、遮挡和指令污染。在LIBERO基准上的大量实验表明,STRONG-VLA持续提升多种VLA架构的任务成功率。在OpenVLA上,面对已见扰动提升达12.60%,未见扰动提升7.77%。类似或更大提升也出现在OpenVLA-OFT(+14.48% / +13.81%)和pi0(+16.49% / +5.58%),证明强跨架构泛化能力。在AIRBOT机器人平台的真实实验进一步验证其实际有效性。结果凸显解耦优化对多模态鲁棒性的关键作用,确立STRONG-VLA为简单而严谨的鲁棒具身控制框架。

原文摘要 · Abstract (English)

Despite their strong performance in embodied tasks, recent Vision-Language-Action (VLA) models remain highly fragile under multimodal perturbations, where visual corruption and linguistic noise jointly induce distribution shifts that degrade task-level execution. Existing robustness approaches typically rely on joint training with perturbed data, treating robustness as a static objective, which leads to conflicting optimization between robustness and task fidelity. In this work, we propose STRONG-VLA, a decoupled fine-tuning framework that explicitly separates robustness acquisition from task-aligned refinement. In Stage I, the model is exposed to a curriculum of multimodal perturbations with increasing difficulty, enabling progressive robustness learning under controlled distribution shifts. In Stage II, the model is re-aligned with clean task distributions to recover execution fidelity while preserving robustness. We further establish a comprehensive benchmark with 28 perturbation types spanning both textual and visual modalities, grounded in realistic sources of sensor noise, occlusion, and instruction corruption. Extensive experiments on the LIBERO benchmark show that STRONG-VLA consistently improves task success rates across multiple VLA architectures. On OpenVLA, our method achieves gains of up to 12.60% under seen perturbations and 7.77% under unseen perturbations. Notably, similar or larger improvements are observed on OpenVLA-OFT (+14.48% / +13.81%) and pi0 (+16.49% / +5.58%), demonstrating strong cross-architecture generalization. Real-world experiments on an AIRBOT robotic platform further validate its practical effectiveness. These results highlight the importance of decoupled optimization for multimodal robustness and establish STRONG-VLA as a simple yet principled framework for robust embodied control.

多模态机器人鲁棒性训练框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。