arXiv:2606.14752cs.CVcs.AI2026-06被引 4

提出X-Tokenizer,让机器人动作具备语义,提升多模态控制精度。

X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining

论文配图:X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
图 1 · 摘自论文原文
  • 用分层残差量化构建带语义的动作编码器,粗粒度表意,细粒度保真。
  • 在240万轨迹上预训练,实现真实世界与仿真环境的顶尖性能。
  • 适合研究视觉-语言-动作联合建模、机器人控制的开发者使用。

现代视觉-语言-动作(VLA)模型需连接预训练的视觉-语言推理与精确的连续机器人控制。现有动作分词方法主要为重建服务,虽保留运动几何但语义监督弱。本文将动作分词视为多模态推理与可执行控制间的语义接口学习,提出轻量级编码器-语义残差量化(SRQ)-解码器架构X-Tokenizer,实现跨多种机械臂形态的统一动作接口。其核心组件SRQ采用非对称结构:第一层通过掩码动作建模(MAM)训练,形成捕捉粗粒度运动意图的离散动作语言;深层仍保持重建导向的残差,保留精细细节。为进一步对齐动作符号与多模态语义,X-Tokenizer在预训练中结合对比对齐预训练基础模型表示空间及下一帧视觉-语言特征预测。在240万条轨迹(20亿动作帧)上预训练后,一个冻结的X-Tokenizer作为混合离散-连续VLA的表征调节信号。该模型在真实世界聚合任务和RoboTwin 2.0仿真中表现领先,相较于FAST在多模态定位任务上提升13.5%,长时程任务提升8.25%,证明动作分词不仅是压缩工具,更是赋能VLA预训练的语义接口。

原文摘要 · Abstract (English)

Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing action tokenizers discretize actions primarily for reconstruction, producing codes that preserve motion geometry but provide only weak semantic supervision to the backbone. We therefore formulate action tokenization not as mere compression, but as semantic interface learning between multimodal reasoning and executable control. To this end, we introduce X-Tokenizer, a lightweight encoder-Semantic Residual Quantization (SRQ)-decoder architecture that provides a shared action interface across diverse robotic arm embodiments. Its key component, SRQ, imposes an asymmetric structure on residual vector quantization: the first level is trained with Masked Action Modeling (MAM) to form a discrete action language that captures coarse motion intent, while deeper levels remain reconstruction-oriented residuals that preserve fine-grained details. To further align action tokens with multimodal semantics, X-Tokenizer is pretrained with contrastive alignment to the representation space of a pretrained foundation model and with next-frame vision-language feature prediction. Pretrained on 2.4M trajectories (2.0B action frames), a single frozen X-Tokenizer plugs into a mixed discrete-continuous VLA as a representation-shaping supervision signal. X-Tokenizer achieves top real-world aggregate and strong RoboTwin 2.0 simulation results. Outperforming FAST in multimodal grounding (+13.5%) and long-horizon tasks (+8.25), it shows that action tokenizers serve as semantic interfaces for VLA pretraining beyond mere action compression.

动作分词多模态机器人控制预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。