arXiv:2507.20186cs.CVeess.IV2025-07中稿 · BMVC 2025被引 3

用小波变换增强图像高频特征,提升SAM模型在复杂任务中的分割效果。

SAMwave: Wavelet-Driven Feature Enrichment for Effective Adaptation of Segment Anything Model

  • 通过小波变换提取多尺度高频特征,替代传统傅里叶方法。
  • 在4个低级视觉任务上显著超越现有适配方法,性能稳定且可解释。
  • 适合需要高精度分割的医学影像、遥感等专业领域应用。

大型基础模型的兴起推动了多个领域的进步。以图像分割为代表的段落一切模型(SAM)展现了显著优势,但其在未训练过的复杂任务上常出现性能下降。现有方法多采用适配器微调并依赖傅里叶域提取高频特征,但分析表明其特征提取能力受限。为此,我们提出 extbf{SAMwave},一种新颖且可解释的方法,利用小波变换从输入数据中提取更丰富的多尺度高频特征。进一步引入复值适配器,通过复小波变换捕捉复值空间-频率信息。通过自适应融合这些小波系数,SAMwave使SAM编码器能捕获更相关的密集预测信息。在四个挑战性低级视觉任务上的实证评估显示,SAMwave显著优于现有适配方法。该性能优势在SAM和SAM2骨干网络上均成立,且适用于真实与复值适配器变体,彰显了方法在适配段落一切模型方面的高效性、灵活性与可解释性。

原文摘要 · Abstract (English)

The emergence of large foundation models has propelled significant advances in various domains. The Segment Anything Model (SAM), a leading model for image segmentation, exemplifies these advances, outperforming traditional methods. However, such foundation models often suffer from performance degradation when applied to complex tasks for which they are not trained. Existing methods typically employ adapter-based fine-tuning strategies to adapt SAM for tasks and leverage high-frequency features extracted from the Fourier domain. However, Our analysis reveals that these approaches offer limited benefits due to constraints in their feature extraction techniques. To overcome this, we propose \textbf{\textit{SAMwave}}, a novel and interpretable approach that utilizes the wavelet transform to extract richer, multi-scale high-frequency features from input data. Extending this, we introduce complex-valued adapters capable of capturing complex-valued spatial-frequency information via complex wavelet transforms. By adaptively integrating these wavelet coefficients, SAMwave enables SAM's encoder to capture information more relevant for dense prediction. Empirical evaluations on four challenging low-level vision tasks demonstrate that SAMwave significantly outperforms existing adaptation methods. This superior performance is consistent across both the SAM and SAM2 backbones and holds for both real and complex-valued adapter variants, highlighting the efficiency, flexibility, and interpretability of our proposed method for adapting segment anything models.

图像分割小波变换模型适配SAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。