arXiv:2412.01034cs.ROcs.CV2024-12被引 25

让机器人控制模型在低资源设备上更快更省电

Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control

  • 训练时模拟低精度量化,提升模型抗误差能力
  • 4位权重量化下实测速度提升2.5倍、能耗降2.5倍
  • 适合边缘计算场景的机器人与自动驾驶部署

基于深度神经网络(DNN)的策略模型,如视觉-语言-动作(VLA)模型,在自动化复杂决策中具有变革性,能处理多模态数据。然而,模型规模扩大导致计算成本激增,给机器人操作和自动驾驶等需快速精准响应的领域带来挑战。为解决资源受限硬件的部署需求,我们提出一种面向模仿学习(IL)策略模型的新量化框架,通过在训练中微调参数以增强对低比特精度误差的鲁棒性,从而在受限条件下保持效率与可靠性。在真实边缘GPU上对机器人操作任务进行4位权重量化评估显示,该框架实现最高2.5倍加速与2.5倍能耗降低;对于4位权重与激活量化的自动驾驶模型,在低端GPU上实现最高3.7倍加速与3.1倍能耗节省。这些结果凸显了该框架在资源受限设备上部署IL策略模型的实际潜力。

原文摘要 · Abstract (English)

Deep neural network (DNN)-based policy models like vision-language-action (VLA) models are transformative in automating complex decision-making across applications by interpreting multi-modal data. However, scaling these models greatly increases computational costs, which presents challenges in fields like robot manipulation and autonomous driving that require quick, accurate responses. To address the need for deployment on resource-limited hardware, we propose a new quantization framework for IL-based policy models that fine-tunes parameters to enhance robustness against low-bit precision errors during training, thereby maintaining efficiency and reliability under constrained conditions. Our evaluations with representative robot manipulation for 4-bit weight-quantization on a real edge GPU demonstrate that our framework achieves up to 2.5x speedup and 2.5x energy savings while preserving accuracy. For 4-bit weight and activation quantized self-driving models, the framework achieves up to 3.7x speedup and 3.1x energy saving on a low-end GPU. These results highlight the practical potential of deploying IL-based policy models on resource-constrained devices.

机器人控制量化边缘计算模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。