融合视觉与文本的轻量级多模态框架,提升海上复杂场景识别准确率。
Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition

- 结合图像、文本描述和分类向量进行多模态融合推理
- 达98%准确率,较SOTA提升3.5%,模型仅68.75MB
- 适合资源受限平台部署,适用于无人艇环境监测
海上多场景识别对智能海洋机器人至关重要,尤其在海洋保护、环境监测和灾害响应中。然而,恶劣海况导致图像质量下降,场景复杂度高,纯视觉模型难以应对。为此,我们提出一种新型多模态AI框架,整合图像数据、文本描述及由多模态大语言模型(MLLM)生成的分类向量,增强语义理解,提升识别精度。采用高效的多模态融合机制,增强模型在复杂海境中的鲁棒性与适应性。实验表明,该模型达到98%准确率,超越先前SOTA模型3.5%。为适配资源受限平台,引入激活感知权重量化(AWQ),将模型压缩至68.75MB,仅损失0.5%准确率,显著降低计算开销。本工作为实时海上场景识别提供高性能解决方案,使自主水面舰艇(ASVs)可在资源受限环境下支持环境监测与应急响应。
原文摘要 · Abstract (English)
Maritime Multi-Scene Recognition is crucial for enhancing the capabilities of intelligent marine robotics, particularly in applications such as marine conservation, environmental monitoring, and disaster response. However, this task presents significant challenges due to environmental interference, where marine conditions degrade image quality, and the complexity of maritime scenes, which requires deeper reasoning for accurate recognition. Pure vision models alone are insufficient to address these issues. To overcome these limitations, we propose a novel multimodal Artificial Intelligence (AI) framework that integrates image data, textual descriptions and classification vectors generated by a Multimodal Large Language Model (MLLM), to provide richer semantic understanding and improve recognition accuracy. Our framework employs an efficient multimodal fusion mechanism to further enhance model robustness and adaptability in complex maritime environments. Experimental results show that our model achieves 98$\%$ accuracy, surpassing previous SOTA models by 3.5$\%$. To optimize deployment on resource-constrained platforms, we adopt activation-aware weight quantization (AWQ) as a lightweight technique, reducing the model size to 68.75MB with only a 0.5$\%$ accuracy drop while significantly lowering computational overhead. This work provides a high-performance solution for real-time maritime scene recognition, enabling Autonomous Surface Vehicles (ASVs) to support environmental monitoring and disaster response in resource-limited settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。