arXiv:2502.03829cs.CV2025-02被引 3

针对手术与深海机器人图像低频信息丢失问题,提出频域增强分割模型

Frequency Domain Enhanced U-Net for Low-Frequency Information-Rich Image Segmentation in Surgical and Deep-Sea Exploration Robots

  • 基于生物视觉机制设计自适应频谱融合模块,平衡跨频段特征
  • 在深海生物与息肉分割任务中达当前最优性能
  • 适合低光照、高噪声环境下需保留低频细节的视觉任务

在深海探测与手术机器人场景中,环境光照与设备分辨率限制常导致高频特征衰减。针对卷积神经网络与人类视觉系统在频带敏感性上的差异(人眼对中频敏感,低频敏感度高于高频),我们实验量化了CNN的对比敏感度函数,并提出受生物视觉启发的波形自适应频谱融合(WASF)方法,以均衡跨频段图像特征。进一步设计感知频率块(PFB),集成WASF以增强频域特征提取能力。基于此构建了FE-UNet模型,采用SAM2骨干网络并融入微调的Hiera-Large模块,在保证分割精度的同时提升泛化能力。实验表明,FE-UNet在海洋生物分割与息肉分割等跨域任务中达到领先水平,展现出强鲁棒性与显著应用潜力。代码将尽快开源。

原文摘要 · Abstract (English)

In deep-sea exploration and surgical robotics scenarios, environmental lighting and device resolution limitations often cause high-frequency feature attenuation. Addressing the differences in frequency band sensitivity between CNNs and the human visual system (mid-frequency sensitivity with low-frequency sensitivity surpassing high-frequency), we experimentally quantified the CNN contrast sensitivity function and proposed a wavelet adaptive spectrum fusion (WASF) method inspired by biological vision mechanisms to balance cross-frequency image features. Furthermore, we designed a perception frequency block (PFB) that integrates WASF to enhance frequency-domain feature extraction. Based on this, we developed the FE-UNet model, which employs a SAM2 backbone network and incorporates fine-tuned Hiera-Large modules to ensure segmentation accuracy while improving generalization capability. Experiments demonstrate that FE-UNet achieves state-of-the-art performance in cross-domain tasks such as marine organism segmentation and polyp segmentation, showcasing robust adaptability and significant application potential. The code will be released soon.

图像分割频域增强机器人视觉医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。