arXiv:2602.15397cs.ROcs.AI2026-02被引 13

提出高效动作分词方法,显著提升视觉语言动作模型性能。

ActionCodec: What Makes for Good Action Tokenizers

  • 基于信息论设计分词原则,优化动作编码的时序重叠与模态关联。
  • 在LIBERO数据集上达95.5%成功率,无需机器人预训练即达新SOTA。
  • 适合关注视觉语言动作模型优化的研究者与开发者使用。

视觉-语言-动作(VLA)模型利用视觉语言模型(VLM)的自回归范式,在指令遵循和训练效率方面表现优异。其核心在于动作分词,但现有设计多聚焦重建保真度,忽视对VLA优化的直接影响。本文从VLA优化角度出发,建立一系列基于信息论的设计原则:最大化时间分词重叠、最小化词表冗余、增强多模态互信息、保证分词独立性。基于此,我们提出高性能动作分词器ActionCodec,显著提升多种仿真与真实场景下的训练效率与模型性能。在LIBERO数据集上,使用ActionCodec微调的SmolVLM2-2.2B模型达到95.5%成功率,无需机器人预训练;经架构增强后达97.4%,刷新无预训练的VLA模型性能纪录。我们相信这些设计原则与开源模型将为社区提供清晰的发展路径。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models leveraging the native autoregressive paradigm of Vision-Language Models (VLMs) have demonstrated superior instruction-following and training efficiency. Central to this paradigm is action tokenization, yet its design has primarily focused on reconstruction fidelity, failing to address its direct impact on VLA optimization. Consequently, the fundamental question of \textit{what makes for good action tokenizers} remains unanswered. In this paper, we bridge this gap by establishing design principles specifically from the perspective of VLA optimization. We identify a set of best practices based on information-theoretic insights, including maximized temporal token overlap, minimized vocabulary redundancy, enhanced multimodal mutual information, and token independence. Guided by these principles, we introduce \textbf{ActionCodec}, a high-performance action tokenizer that significantly enhances both training efficiency and VLA performance across diverse simulation and real-world benchmarks. Notably, on LIBERO, a SmolVLM2-2.2B fine-tuned with ActionCodec achieves a 95.5\% success rate without any robotics pre-training. With advanced architectural enhancements, this reaches 97.4\%, representing a new SOTA for VLA models without robotics pre-training. We believe our established design principles, alongside the released model, will provide a clear roadmap for the community to develop more effective action tokenizers.

动作分词VLA模型自回归多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。