重新审视机器人视觉语言动作模型的规模化,发现数据混合与对齐方式至关重要。
Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
- 采用统一末端执行器相对动作表示,提升跨机器人迁移能力。
- 盲目合并异构机器人数据反而导致性能下降,存在负向迁移风险。
- 常规正则化策略在大规模训练中效果不显著,需重新评估其有效性。
尽管视觉-语言-动作(VLA)模型在通用机器人控制方面展现出巨大潜力,但其是否以及在何种条件下可沿用标准“扩大数据量”的方法仍不明确,尤其在不同机器人形态、传感器和动作空间下数据本就异质的情况下。本文通过系统性、受控的实验研究,重新评估了跨多种机器人的预训练核心设计选择。基于一个结合视觉-语言主干与流匹配的代表性VLA框架,在模拟和真实机器人实验中对比关键决策。为提高真实世界结果的可靠性,引入分组盲法集成协议,使操作者无法识别模型身份,并将策略执行与结果评估分离,减少人为偏见。分析聚焦三个维度:(1) 物理对齐:发现统一末端执行器(EEF)相对动作表示对跨形态迁移至关重要;(2) 机器人混合:发现简单拼接异构机器人数据集常引发负向迁移而非增益,凸显无差别数据扩增的脆弱性;(3) 训练正则化:观察到如感官丢弃和多阶段微调等直观策略在规模化时并不一致提升性能。该研究挑战了部分关于具身化扩展的常见假设,为从多样化机器人数据中训练大规模VLA策略提供了实用指导。
原文摘要 · Abstract (English)
While Vision-Language-Action (VLA) models show strong promise for generalist robot control, it remains unclear whether -- and under what conditions -- the standard "scale data" recipe translates to robotics, where training data is inherently heterogeneous across embodiments, sensors, and action spaces. We present a systematic, controlled study of VLA scaling that revisits core training choices for pretraining across diverse robots. Using a representative VLA framework that combines a vision-language backbone with flow-matching, we ablate key design decisions under matched conditions and evaluate in extensive simulation and real-robot experiments. To improve the reliability of real-world results, we introduce a Grouped Blind Ensemble protocol that blinds operators to model identity and separates policy execution from outcome judgment, reducing experimenter bias. Our analysis targets three dimensions of VLA scaling. (1) Physical alignment: we show that a unified end-effector (EEF)-relative action representation is critical for robust cross-embodiment transfer. (2) Embodiment mixture: we find that naively pooling heterogeneous robot datasets often induces negative transfer rather than gains, underscoring the fragility of indiscriminate data scaling. (3) Training regularization: we observe that intuitive strategies, such as sensory dropout and multi-stage fine-tuning, do not consistently improve performance at scale. Together, this study challenge some common assumptions about embodied scaling and provide practical guidance for training large-scale VLA policies from diverse robotic data. Project website: https://research.beingbeyond.com/rethink_vla
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。