arXiv:2603.15374cs.CV2026-03

通过频域修正提升医学影像深度估计,参数高效且效果顶尖

Spectral Rectification for Parameter-Efficient Adaptation of Foundation Models in Colonoscopy Depth Estimation

  • 用可学习小波分解增强图像高频特征,精准修复模型误判根源
  • 在C3VD和SimCol3D上相对误差低至0.022和0.027,性能领先
  • 适合医疗视觉任务中需轻量适配大模型的场景

准确的单目深度估计对结肠镜检查中的病灶定位与导航至关重要。在自然图像上训练的基础模型无法直接泛化到结肠镜图像。我们发现核心问题并非语义差异,而是频域上的统计偏移:结肠镜图像缺乏模型依赖的强高频边缘与纹理梯度。为此,提出SpecDepth框架,以参数高效方式保留预训练模型的几何表征能力,同时适配结肠镜域。其关键创新是自适应频谱校正模块,通过可学习小波分解显式建模并放大特征图中衰减的高频成分。不同于传统微调可能破坏高层语义特征,该方法仅进行低层针对性调整,使输入信号重新匹配原始模型的归纳偏置。在公开数据集C3VD和SimCol3D上,该方法分别达到0.022和0.027的绝对相对误差,性能达当前最优。结果表明,直接解决频谱失配是适配视觉基础模型至专业医学影像任务的有效策略。代码将在论文被接收后公开。

原文摘要 · Abstract (English)

Accurate monocular depth estimation is critical in colonoscopy for lesion localization and navigation. Foundation models trained on natural images fail to generalize directly to colonoscopy. We identify the core issue not as a semantic gap, but as a statistical shift in the frequency domain: colonoscopy images lack the strong high-frequency edge and texture gradients that these models rely on for geometric reasoning. To address this, we propose SpecDepth, a parameter-efficient adaptation framework that preserves the robust geometric representations of the pre-trained models while adapting to the colonoscopy domain. Its key innovation is an adaptive spectral rectification module, which uses a learnable wavelet decomposition to explicitly model and amplify the attenuated high-frequency components in feature maps. Different from conventional fine-tuning that risks distorting high-level semantic features, this targeted, low-level adjustment realigns the input signal with the original inductive bias of the foundational model. On the public C3VD and SimCol3D datasets, SpecDepth achieved state-of-the-art performance with an absolute relative error of 0.022 and 0.027, respectively. Our work demonstrates that directly addressing spectral mismatches is a highly effective strategy for adapting vision foundation models to specialized medical imaging tasks. The code will be released publicly after the manuscript is accepted for publication.

深度估计医学影像参数高效频域处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。