让小型视觉语言模型在边缘机器人上实时完成感知与移动决策
Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots
- 将轻量级视觉语言模型嵌入机器人本地硬件,实现边端协同推理
- 在仅用机载设备条件下完成动态环境下的实时感知与移动
- 适合资源受限场景,如救灾、服务机器人和军事应用
在无GPS环境中,自主机器人依赖本地、高效推理能力。本文展示了在移动机器人上部署小型视觉语言模型(VLM)的可行性,使其在严格计算约束下实现实时场景理解与推理。与以往将感知与运动分离的方法不同,该框架利用机载硬件实现动态环境中的同步移动与推理。系统集成紧凑型VLM与多模态感知,直接在嵌入式设备上完成上下文解释,无需云端支持。实验验证了计算效率、任务准确率与系统响应间的平衡。在移动机器人上的实现,是首个成功将小型VLM用于边缘端并发推理与移动的案例。本工作为服务机器人、灾难响应和国防应用中的可扩展、可信自主系统奠定了基础。
原文摘要 · Abstract (English)
The deployment of artificial intelligence models at the edge is increasingly critical for autonomous robots operating in GPS-denied environments where local, resource-efficient reasoning is essential. This work demonstrates the feasibility of deploying small Vision-Language Models (VLMs) on mobile robots to achieve real-time scene understanding and reasoning under strict computational constraints. Unlike prior approaches that separate perception from mobility, the proposed framework enables simultaneous movement and reasoning in dynamic environments using only on-board hardware. The system integrates a compact VLM with multimodal perception to perform contextual interpretation directly on embedded hardware, eliminating reliance on cloud connectivity. Experimental validation highlights the balance between computational efficiency, task accuracy, and system responsiveness. Implementation on a mobile robot confirms one of the first successful deployments of small VLMs for concurrent reasoning and mobility at the edge. This work establishes a foundation for scalable, assured autonomy in applications such as service robotics, disaster response, and defense operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。