arXiv:2510.25122cs.RO2025-10被引 13

轻量化视觉语言模型让机器人在低算力设备上实现高效精准操作

NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies

  • 视觉语言解耦,延迟和计算开销更低
  • 边缘设备推理速度提升52倍,参数量减少98%
  • 动态路由适配任务复杂度,适合资源受限的机器人部署

视觉-语言-动作(VLA)模型通过整合视觉-语言模型(VLM)与动作解码器,显著推动了机器人操作的发展。然而,在移动机器人或嵌入式系统(如Jetson Orin Nano)等资源受限的边缘设备上部署仍面临高计算需求的挑战,尤其在功耗、延迟和算力受限的真实场景中。为此,我们提出纳米级视觉-语言-动作模型(NanoVLA),一组轻量级VLA架构,在极低资源下实现高性能。核心创新包括:(1)将传统早期视觉与语言融合改为晚期融合,提升性能并支持缓存,降低推理开销与延迟;(2)长短动作块分段策略,确保多步规划流畅连贯且不失实时响应;(3)动态路由机制,根据任务复杂度自适应选择轻量或重型骨干网络,进一步优化推理效率。在多个基准测试及真实场景部署中,NanoVLA相比现有最先进VLA模型,在边缘设备上实现最高52倍的推理加速,参数量减少98%,同时保持或超越其任务准确率与泛化能力。消融实验验证了解耦策略维持跨任务迁移性,路由模块显著提升性价比,使高精度机器人操作在资源受限硬件上成为可能。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have significantly advanced robotic manipulation by integrating vision-language models (VLMs), and action decoders into a unified architecture. However, their deployment on resource-constrained edge devices, such as mobile robots or embedded systems (e.g., Jetson Orin Nano), remains challenging due to high computational demands, especially in real-world scenarios where power, latency, and computational resources are critical. To close this gap, we introduce Nano-scale Vision-Language Action (NanoVLA), a family of lightweight VLA architectures that achieve high performance with minimal resources. Our core innovations include: (1) vision-language decoupling that moves conventional early vision and language inputs fusion in VLM to late stage, achieving better performance while enabling caching and reduce inference overhead and latency; (2) long-short action chunking to ensure smooth, coherent multi-step planning without sacrificing real-time responsiveness; (3) dynamic routing that adaptively assigns lightweight or heavy backbones based on task complexity, further optimizing inference efficiency. Experimental results on several benchmarks, as well as real-world deployments, demonstrate that NanoVLA achieves up to 52x faster inference on edge devices compared to previous state-of-the-art VLA models, with 98% less parameters while maintaining or surpassing their task accuracy and generalization. Ablation studies confirm that our decoupling strategy preserves cross-task transferability, and the routing module enhances cost-performance trade-offs, enabling practical, high-precision robotic manipulation on resource-constrained hardware.

机器人轻量化边缘计算视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。