用海量数据训练的向量量化动作分词器,让机器人动作更顺滑、推理更快。
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

- 基于超大规模动作轨迹数据,用向量量化构建动作分词器。
- 合成数据越多,长程任务成功率越高,实测提升达30%。
- 零样本适配多种任务,适合实时机器人控制应用。
本文提出一种基于向量量化的动作分词器,依托迄今最大规模的动作轨迹数据集,数据量比以往方法多100倍以上。该数据集使分词器能捕捉丰富的时空动态,显著加速推理并生成更平滑、连贯的动作输出。训练完成后,分词器可零样本适配多种下游任务,涵盖短时反应行为与长时规划。关键发现是:合成与真实动作轨迹之间的领域差异微小,因此可大量使用合成数据训练而不影响真实表现。我们在仿真环境和真实机器人平台上进行了广泛实验,结果表明,随着合成轨迹数据量增加,下游任务性能显著提升,尤其在两个长程任务中成功率达30%以上。这证明该分词器是实现高效可靠具身智能系统的有力方案。
原文摘要 · Abstract (English)
In this paper, we introduce an innovative vector quantization based action tokenizer built upon the largest-scale action trajectory dataset to date, leveraging over 100 times more data than previous approaches. This extensive dataset enables our tokenizer to capture rich spatiotemporal dynamics, resulting in a model that not only accelerates inference but also generates smoother and more coherent action outputs. Once trained, the tokenizer can be seamlessly adapted to a wide range of downstream tasks in a zero-shot manner, from short-horizon reactive behaviors to long-horizon planning. A key finding of our work is that the domain gap between synthetic and real action trajectories is marginal, allowing us to effectively utilize a vast amount of synthetic data during training without compromising real-world performance. To validate our approach, we conducted extensive experiments in both simulated environments and on real robotic platforms. The results demonstrate that as the volume of synthetic trajectory data increases, the performance of our tokenizer on downstream tasks improves significantly-most notably, achieving up to a 30% higher success rate on two real-world tasks in long-horizon scenarios. These findings highlight the potential of our action tokenizer as a robust and scalable solution for real-time embodied intelligence systems, paving the way for more efficient and reliable robotic control in diverse application domains.Project website: https://xiaoxiao0406.github.io/vqvla.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。