评估视觉语言模型在水下机器人感知中的表现与不确定性
Assessing Vision-Language Models for Perception in Autonomous Underwater Robotic Software
- 用实证方法测试VLM在水下环境的感知能力
- 发现模型对水下垃圾检测性能差异显著,且不确定性与性能相关
- 为水下机器人开发提供选型依据,适合工程团队参考
自主水下机器人(AUR)在低能见度和恶劣水况中运行,给感知模块开发带来挑战。尽管深度学习被用于支持其运作,但水下数据稀缺且噪声大,影响模型可信度。视觉语言模型(VLM)可通过上下文线索泛化到未见物体,在噪声环境中保持鲁棒性,具备应用潜力。然而,其在水下环境下的性能与不确定性尚未从软件工程角度充分研究。基于工业伙伴对海事系统可靠性与风险管控的需求,本文开展实证评估,考察VLM在水下垃圾检测任务中的表现、不确定性及其关系,帮助软件工程师选择合适的VLM用于AUR软件。
原文摘要 · Abstract (English)
Autonomous Underwater Robots (AURs) operate in challenging underwater environments, including low visibility and harsh water conditions. Such conditions present challenges for software engineers developing perception modules for the AUR software. To successfully carry out these tasks, deep learning has been incorporated into the AUR software to support its operations. However, the unique challenges of underwater environments pose difficulties for deep learning models, which often rely on labeled data that is scarce and noisy. This may undermine the trustworthiness of AUR software that relies on perception modules. Vision-Language Models (VLMs) offer promising solutions for AUR software as they generalize to unseen objects and remain robust in noisy conditions by inferring information from contextual cues. Despite this potential, their performance and uncertainty in underwater environments remain understudied from a software engineering perspective. Motivated by the needs of an industrial partner in assurance and risk management for maritime systems to assess the potential use of VLMs in this context, we present an empirical evaluation of VLM-based perception modules within the AUR software. We assess their ability to detect underwater trash by computing performance, uncertainty, and their relationship, to enable software engineers to select appropriate VLMs for their AUR software.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。