通过多重一致性约束,提升视觉语言动作模型在复杂环境下的鲁棒性。
RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

- 引入指令、轨迹、观测三重一致性约束,增强模型稳定性。
- 在LIBERO-Plus等数据集上性能超越基线,抗干扰能力显著提升。
- 适合关注具身智能鲁棒控制的研究者与开发者。
视觉语言动作(VLA)模型在具身操作任务中表现强劲,但在视觉变化、语言指令改写及多重扰动下仍显脆弱。这表明现有方法仍依赖训练分布中的浅层关联,而非学习任务语义、环境状态与动作生成之间的稳定映射。尽管近期工作通过大规模训练、后训练微调或增强预测建模提升了鲁棒性,但极少在端到端策略中强制施加不变性一致性。为此,本文提出RoVLA框架,引入多一致性约束:指令一致性(IC)确保语义等价指令下的稳定语义对齐;演化一致性(EC)保持动作意图在生成过程中的连贯性;观测一致性(OC)通过在目标扰动前后强制一致预测,提升对视觉与本体感知扰动的鲁棒性。通过显式建模这些不变性,RoVLA降低对表层相关性的依赖,增强泛化能力。在LIBERO-Plus、RoboTwin 2.0及真实世界操作任务上的实验表明,RoVLA持续优于强基线方法,在多种任务与观测变化下表现出更优鲁棒性。结果验证了多一致性学习在鲁棒具身控制中的有效性。代码将公开于https://github.com/HCPLab-SYSU/RoVLA。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have shown strong performance on embodied manipulation, yet they remain brittle under visual observation changes, paraphrased language instructions, and compounded perturbations. This limitation suggests that existing methods still rely heavily on shallow correlations in the training distribution, rather than learning stable couplings among task semantics, environment states, and action generation. Although recent efforts improve robustness through larger-scale training, post-training adaptation, or enhanced predictive modeling, they rarely enforce invariance-oriented consistency within the end-to-end policy itself. To address this issue, we propose RoVLA, a robust vision-language-action framework with multi-consistency constraints. RoVLA enforces consistency under three complementary transformations: instruction semantics, trajectory evolution, and observation perturbation. Specifically, Instructional Consistency (IC) promotes stable grounding under semantically equivalent instruction rewrites, Evolutionary Consistency (EC) preserves coherent action intent throughout the generation process, and Observational Consistency (OC) improves robustness to visual and proprioceptive perturbations by enforcing consistent predictions before and after targeted disturbances. By explicitly modeling these invariances during training, RoVLA reduces reliance on superficial correlations and improves robustness and generalization. Experiments on LIBERO-Plus, RoboTwin 2.0, and real-world manipulation tasks show that RoVLA consistently outperforms strong baseline methods and exhibits superior robustness under diverse task and observation shifts. These results demonstrate the effectiveness of multi-consistency learning for robust embodied control. Codes will be available at https://github.com/HCPLab-SYSU/RoVLA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。