用物理模型增强自蒸馏,让声呐图像识别更抗视角干扰。
BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations

- 基于物理规律设计噪声和衰减模拟,强制模型忽略视角变化。
- 仅用10%标注数据就达71.4% mIoU,性能接近全量数据的96%。
- 适合水下地形分析、无人潜航器感知等需要鲁棒特征的任务。
侧扫声呐图像的自动感知严重受限于物理声学伪影,导致表征同时包含海底固有反射率与瞬时观测几何。现有自监督学习框架依赖自然图像的增强方式,未考虑声学退化且未显式强制视图不变性。为此,我们提出一种基于DINOv3架构、采用ConvNeXt-v2-Tiny主干的物理信息自蒸馏框架,以最大化数据效率。该方法通过两类机制实现视图不变性:一是模拟斑点噪声、距离相关衰减和辐射校准偏差的物理增强;二是利用希尔伯特-施密特独立性准则(HSIC)惩罚项,显式解耦密集补丁特征与物理观测参数。此外,提出跨四个网络阶段的密集分层特征融合策略,兼顾细粒度沉积物细节与深层语义抽象。大量实验表明,该框架无需人工标注即可将复杂海底地形稳定聚类为无噪声语义簇。在S3Seg数据集的监督下游任务中,融合表示表现出卓越的数据效率:仅使用10%标注数据即达到绝对峰值性能的96%,最终实现71.4%的平均交并比(mIoU)和86.5%的整体准确率。
原文摘要 · Abstract (English)
Automated perception in side-scan sonar (SSS) imagery is severely hindered by physical acoustic artifacts, resulting in representations that inextricably mix intrinsic seabed reflectivity with transient viewing geometries. Existing self-supervised learning (SSL) frameworks rely on augmentations designed for natural images, failing to account for acoustic degradation and explicitly enforce view-invariance. To address this gap, we introduce a physics-informed self-distillation framework built upon the DINOv3 architecture utilizing a ConvNeXt-v2-Tiny backbone to maximize data efficiency. The proposed methodology enforces view-invariance through two primary mechanisms: physically motivated augmentations that simulate speckle noise, range-dependent attenuation, and radiometric miscalibration; and a Hilbert-Schmidt Independence Criterion (HSIC) penalty that explicitly decouples learned dense patch features from physical viewing parameters. Furthermore, we propose a dense, hierarchical feature fusion strategy across all four network stages to preserve fine-grained sediment details alongside deep semantic abstractions. Extensive evaluation demonstrates that the framework natively groups complex benthic topographies into stable, noise-free semantic clusters without relying on manual annotations. During supervised downstream tasks on the S3Seg dataset, the fused representations exhibited exceptional data efficiency, achieving 96% of its absolute peak performance using only 10% of the available annotated data, ultimately reaching a mean Intersection over Union (mIoU) of 71.4% and an overall accuracy of 86.5%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。