用手机AI自动发送水下情境消息,提升潜水安全与沟通效率。
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
- 基于移动端视觉语言模型,自动解析水下场景生成消息
- 通过分层生成与抗错训练,提升通信在水下环境中的可靠性
- 支持真实潜水场景,适合个人潜水员及移动多模态应用
潜水活动每年吸引数百万人探索海洋,但保持情境感知和有效沟通对安全至关重要。传统水下通信设备笨重昂贵,难以普及;现有轻量级系统依赖预设文本,无法适应具体场景。本文提出AquaVLM,一种基于智能手机的“点击发送”式水下通信系统,利用在自动生成的水下对话数据集上微调的移动端视觉语言模型(VLM),结合分层消息生成流程,实现上下文感知的消息自动创建。系统协同设计了VLM与传输机制,引入抗错微调以增强对传输错误的鲁棒性。我们开发了虚拟现实模拟器,用于在逼真水下环境中体验系统,并在iOS平台实现了可运行原型。主观与客观评估均验证了AquaVLM的有效性,展示了其在个人潜水通信及更广泛移动VLM应用中的潜力。
原文摘要 · Abstract (English)
Underwater activities like scuba diving enable millions annually to explore marine environments for recreation and scientific research. Maintaining situational awareness and effective communication are essential for diver safety. Traditional underwater communication systems are often bulky and expensive, limiting their accessibility to divers of all levels. While recent systems leverage lightweight smartphones and support text messaging, the messages are predefined and thus restrict context-specific communication. In this paper, we present AquaVLM, a tap-and-send underwater communication system that automatically generates context-aware messages and transmits them using ubiquitous smartphones. Our system features a mobile vision-language model (VLM) fine-tuned on an auto-generated underwater conversation dataset and employs a hierarchical message generation pipeline. We co-design the VLM and transmission, incorporating error-resilient fine-tuning to improve the system's robustness to transmission errors. We develop a VR simulator to enable users to experience AquaVLM in a realistic underwater environment and create a fully functional prototype on the iOS platform for real-world experiments. Both subjective and objective evaluations validate the effectiveness of AquaVLM and highlight its potential for personal underwater communication as well as broader mobile VLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。