自监督压缩与去噪一体框架,让水下声呐实时传输更清晰
Self-Supervised Compression and Artifact Correction for Streaming Underwater Imaging Sonar
- 通过频率编码潜变量与多尺度分割,联合压缩与消除斑点等伪影
- 在0.0118 bpp下SSIM达0.77,比现有方法提升40%,带宽降低超80%
- 已部署于三处河流实测,支持鲑鱼计数与环境监测
实时声呐成像对水下监测至关重要,但受限于低上行带宽及严重声呐特有伪影(斑点、运动模糊、混响、声阴影),高达98%的帧受影响。我们提出SCOPE,一种无需干净-噪声配对或合成假设的自监督框架,联合执行压缩与去噪。该框架结合(i)自适应码本压缩(ACC),学习针对声呐的频域编码潜变量;(ii)频域感知多尺度分割(FAMS),将图像分解为低频结构与稀疏高频动态,同时抑制快速波动伪影。通过无标签生成的低通代理对,采用规避式训练策略引导频域学习。在数月实地ARIS声呐数据上评估,SCOPE在比特率低至≤0.0118 bpp时达到0.77的结构相似性指数(SSIM),较先前自监督去噪基线提升40%,上行带宽减少超80%,并提升下游检测性能。系统实现实时运行:嵌入式GPU编码仅需3.1毫秒,服务端全层解码耗时97毫秒。该系统已在太平洋西北部三条河流部署数月,支持实时鲑鱼计数与野外环境监测。结果表明,学习频域结构化潜变量可在真实场景下实现低比特率、保细节的声呐流传输。
原文摘要 · Abstract (English)
Real-time imaging sonar is crucial for underwater monitoring where optical sensing fails, but its use is limited by low uplink bandwidth and severe sonar-specific artifacts (speckle, motion blur, reverberation, acoustic shadows) affecting up to 98% of frames. We present SCOPE, a self-supervised framework that jointly performs compression and artifact correction without clean-noise pairs or synthetic assumptions. SCOPE combines (i) Adaptive Codebook Compression (ACC), which learns frequency-encoded latent representations tailored to sonar, with (ii) Frequency-Aware Multiscale Segmentation (FAMS), which decomposes frames into low-frequency structure and sparse high-frequency dynamics while suppressing rapidly fluctuating artifacts. A hedging training strategy further guides frequency-aware learning using low-pass proxy pairs generated without labels. Evaluated on months of in-situ ARIS sonar data, SCOPE achieves a structural similarity index (SSIM) of 0.77, representing a 40% improvement over prior self-supervised denoising baselines, at bitrates down to <= 0.0118 bpp. It reduces uplink bandwidth by more than 80% while improving downstream detection. The system runs in real time, with 3.1 ms encoding on an embedded GPU and 97 ms full multi-layer decoding on the server end. SCOPE has been deployed for months in three Pacific Northwest rivers to support real-time salmon enumeration and environmental monitoring in the wild. Results demonstrate that learning frequency-structured latents enables practical, low-bitrate sonar streaming with preserved signal details under real-world deployment conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。