一个模型实现声呐与光学图像双向转换,性能接近专用模型。
One Model, Two Worlds: Bidirectional Sonar-Optical Translation

- 用不对称路径分别处理声呐与光学的成像物理差异。
- 双向转换中声呐转光学的PSNR仅差0.11 dB,光学转声呐的FID提升0.70。
- 自适应感知监督机制减少训练成本,适合水下视觉系统集成。
声呐与光学图像之间的相互转换对水下感知至关重要,但传统方法需分别部署模型,造成存储与计算冗余。本文提出统一的双向模型,突破对称设计局限:引入方向非对称现实桥(DARB),共享扩散桥主干,通过范围感知调制(声呐→光学)和极角射线依赖处理(光学→声呐)分别建模不同成像物理。实验表明,对称训练会降低声呐→光学的PSNR达2.60 dB。为此设计自适应现实监督(ARS),依据重建质量与梯度平衡动态调节感知监督强度。最终,该模型在声呐→光学上比专用模型低0.11 dB PSNR,优于光学→声呐专用模型0.70 FID,且在八项指标中超越两个独立训练的双向扩散模型(BBDMs)中的七项。
原文摘要 · Abstract (English)
Translating between imaging sonar and optical cameras is valuable for underwater perception, but supporting both directions with separate models duplicates storage and computation. A unified bidirectional model is therefore attractive, yet existing approaches largely treat the two directions symmetrically despite their fundamentally different image-formation physics. We argue that sharing a generative model does not require sharing the physics. We introduce the Direction-Asymmetric Realism Bridge (DARB), which retains a shared diffusion-bridge trunk while routing direction-specific physical priors through asymmetric pathways: range-aware modulation for sonar-to-optical translation and polar ray-dependent processing for optical-to-sonar translation. We further show that symmetry in training is also costly: applying a common realism schedule reduces sonar-to-optical PSNR by 2.60 dB. Our Adaptive Realism Supervision (ARS) instead determines when, where, and how strongly perceptual supervision is applied from reconstruction quality and gradient balance. Together, DARB and ARS enable one bidirectional model to match the sonar-to-optical specialist within 0.11 dB PSNR, outperform the optical-to-sonar specialist by 0.70 FID, and surpass two independently trained BBDMs on seven of eight metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。