arXiv:2511.02427cs.CVcs.RO2025-11中稿 · AI-2025 Forty-fift…被引 1

测试小模型在边缘设备上零样本场景理解的可行性与局限

From the Laboratory to Real-World Application: Evaluating Zero-Shot Scene Interpretation on Edge Devices for Mobile Robotics

  • 用轻量级视觉语言模型实现边缘端零样本场景解析
  • 实测在真实城市、校园和室内场景中准确率超60%
  • 适合移动机器人部署,揭示模型固有偏见与实用瓶颈

视频理解、场景解析与常识推理是使智能体感知环境并作出合理决策的关键挑战。近年来,大型语言模型(LLMs)和视觉语言模型(VLMs)在这些领域取得显著进展,支持特定领域应用及零样本开放词汇任务。然而,其高计算复杂度限制了在边缘设备和移动机器人中的应用,尤其在精度与推理时延间存在权衡。本文研究前沿小规模VLM在场景解析与动作识别任务上的表现,重点评估其在移动机器人背景下部署于边缘设备的可行性。实验在包含真实城市景观、校园及室内场景的多样化数据集上进行,分析模型在实际环境中的潜力、挑战、弱点及固有偏差,并探讨所获信息的应用价值。

原文摘要 · Abstract (English)

Video Understanding, Scene Interpretation and Commonsense Reasoning are highly challenging tasks enabling the interpretation of visual information, allowing agents to perceive, interact with and make rational decisions in its environment. Large Language Models (LLMs) and Visual Language Models (VLMs) have shown remarkable advancements in these areas in recent years, enabling domain-specific applications as well as zero-shot open vocabulary tasks, combining multiple domains. However, the required computational complexity poses challenges for their application on edge devices and in the context of Mobile Robotics, especially considering the trade-off between accuracy and inference time. In this paper, we investigate the capabilities of state-of-the-art VLMs for the task of Scene Interpretation and Action Recognition, with special regard to small VLMs capable of being deployed to edge devices in the context of Mobile Robotics. The proposed pipeline is evaluated on a diverse dataset consisting of various real-world cityscape, on-campus and indoor scenarios. The experimental evaluation discusses the potential of these small models on edge devices, with particular emphasis on challenges, weaknesses, inherent model biases and the application of the gained information. Supplementary material is provided via the following repository: https://datahub.rz.rptu.de/hstr-csrl-public/publications/scene-interpretation-on-edge-devices/

边缘计算视觉语言模型移动机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。