针对视觉语言动作模型量化后轨迹漂移问题,提出感知漂移的量化方法。
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
- 将量化建模为序列决策中的漂移感知优化问题。
- 在低比特下显著降低轨迹漂移,性能接近全精度模型。
- 适合资源受限机器人上部署视觉语言动作模型。
视觉语言动作模型(VLAs)在具身智能中展现巨大潜力,但其在资源受限机器人上的部署仍面临高内存与计算需求的挑战。虽然训练后量化(PTQ)提供了高效解决方案,但直接应用于VLAs常导致序列控制时性能严重下降。我们发现,时间误差累积是关键因素:在视觉-语言到动作接口处的量化扰动被逐步放大,引发执行轨迹的运动学漂移。为此,提出感知漂移的训练后量化(DA-PTQ),将量化建模为序列决策过程中的漂移感知优化问题。DA-PTQ包含两部分:(1) 跨空间表示补偿,缓解多模态表示与动作空间间的结构失真,提升动作一致性;(2) 基于运动的混合精度分配,通过最小化轨迹级运动误差来分配比特位数。大量实验表明,DA-PTQ显著减少运动学漂移,在低比特设置下性能接近全精度模型,使VLAs可在资源受限机器人平台上实现实用部署。
原文摘要 · Abstract (English)
Vision-Language-Action models (VLAs) have demonstrated strong potential for embodied AI, yet their deployment on resource-limited robots remains challenging due to high memory and computational demands. While Post-Training Quantization (PTQ) provides an efficient solution, directly applying PTQ to VLAs often results in severe performance degradation during sequential control. We identify temporal error accumulation as a key factor, where quantization perturbations at the vision-language-to-action interface are progressively amplified, leading to kinematic drift in executed trajectories. To address this issue, we propose Drift-Aware Post-Training Quantization (DA-PTQ), which formulates quantization as a drift-aware optimization problem over sequential decision processes. DA-PTQ consists of two components: (1) Cross-Space Representation Compensation, which mitigates structured distortions between multimodal representations and action space to improve action consistency, and (2) Motion-Driven Mixed-Precision Allocation, which assigns bit-widths by minimizing trajectory-level motion errors. Extensive experiments show that DA-PTQ significantly reduces kinematic drift and achieves comparable performance to full-precision models under low-bit settings, enabling practical deployment of VLAs on resource-limited robotic platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。