arXiv:2512.04952cs.CVcs.RO2025-12被引 18

用可学习的动作标记化提升视觉语言动作模型效率与性能

FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization

  • 设计可学习的动作标记器,将动作块编码为单通道图像,压缩比高且保留时空依赖
  • 采用分块自回归解码和轻量专家模型,推理速度更快,任务表现超越现有最佳模型
  • 在模拟与真实场景中均展现强泛化能力,适合需要高效推理的机器人学习任务

自回归视觉语言动作(VLA)模型在机器人操作任务中表现出强大能力,但其核心动作标记化过程常面临重建精度与推理效率的权衡。本文提出FASTer框架,通过可学习标记器与自回归策略的结合,实现高效通用的机器人学习。FASTerVQ将动作块编码为单通道图像,捕捉全局时空依赖,同时保持高压缩比;FASTerVLA在此基础上采用分块自回归解码与轻量动作专家,显著提升推理速度与任务性能。在多种模拟与真实世界基准测试中,FASTerVQ实现更优重建质量、高标记利用率及跨任务、跨硬件平台的强泛化能力;FASTerVLA进一步超越现有最优VLA模型,在推理速度与任务表现上均取得突破。

原文摘要 · Abstract (English)

Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and inference efficiency. We introduce FASTer, a unified framework for efficient and generalizable robot learning that integrates a learnable tokenizer with an autoregressive policy built upon it. FASTerVQ encodes action chunks as single-channel images, capturing global spatio-temporal dependencies while maintaining a high compression ratio. FASTerVLA builds on this tokenizer with block-wise autoregressive decoding and a lightweight action expert, achieving both faster inference and higher task performance. Extensive experiments across simulated and real-world benchmarks show that FASTerVQ delivers superior reconstruction quality, high token utilization, and strong cross-task and cross-embodiment generalization, while FASTerVLA further improves overall capability, surpassing previous state-of-the-art VLA models in both inference speed and task performance.

机器人学习自回归模型动作标记化高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。