MERLION让水下机器人智能挑选关键画面并增强模糊影像。
MERLION: Marine ExploRation with Language guIded Online iNformative Visual Sampling and Enhancement
- 用图文模型对齐用户需求与视觉样本
- 在浑浊水域中提升图像清晰度,提升信息量
- 适合水下监测、生态研究等需要高效采样的场景
基于自主水下航行器(AUV)的自主化、目标导向型水下视觉监测与探索面临在线和离线双重挑战。在线约束包括有限的机载存储与通信带宽,离线约束则在于从视频数据中人工筛选关键帧耗时费力。例如,在长序列的水下观测中定位最具信息量的鱼类画面。该问题在能见度低的浑浊水域中尤为严峻。本文提出MERLION框架,实现语义对齐且视觉增强的水下海洋环境监控摘要。其整合三部分:(a) 图文模型实现视觉样本与用户需求的语义对齐;(b) 图像增强模型处理浑浊水下视觉数据;(c) 信息采样器生成监控经验摘要。我们在真实数据上通过用户研究验证了该框架,并采用评估指标进行定性与定量分析,结果优于现有最先进方法。代码已开源:https://github.com/MARVL-Lab/MERLION.git。
原文摘要 · Abstract (English)
Autonomous and targeted underwater visual monitoring and exploration using Autonomous Underwater Vehicles (AUVs) can be a challenging task due to both online and offline constraints. The online constraints comprise limited onboard storage capacity and communication bandwidth to the surface, whereas the offline constraints entail the time and effort required for the selection of desired key frames from the video data. An example use case of targeted underwater visual monitoring is finding the most interesting visual frames of fish in a long sequence of an AUV's visual experience. This challenge of targeted informative sampling is further aggravated in murky waters with poor visibility. In this paper, we present MERLION, a novel framework that provides semantically aligned and visually enhanced summaries for murky underwater marine environment monitoring and exploration. Specifically, our framework integrates (a) an image-text model for semantically aligning the visual samples to the users' needs, (b) an image enhancement model for murky water visual data and (c) an informative sampler for summarizing the monitoring experience. We validate our proposed MERLION framework on real-world data with user studies and present qualitative and quantitative results using our evaluation metric and show improved results compared to the state-of-the-art approaches. We have open-sourced the code for MERLION at the following link https://github.com/MARVL-Lab/MERLION.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。