arXiv:2607.26395cs.CV2026-07

解决内镜图像配准不准导致的分割误差问题。

Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation

  • 通过可靠度引导的复数域融合,区分WLI和NBI不同作用
  • 在真实数据集上提升病灶分割精度,边界更清晰
  • 适合需要高精度内镜图像分析的研究者

白光成像(WLI)和窄带成像(NBI)能互补提供内镜病灶信息,但因视角变化、组织形变及手持连续采集,常出现空间错位。直接融合易混合非对应区域,甚至恶化病灶边界分割。为此,提出一种可靠性感知的复数域融合框架,先建立拓扑正则化的特征对应关系,再估计跨模态对应关系的可靠性。基于此可靠性,模型在可学习的复数表示中选择性融合WLI与NBI特征:WLI主要贡献外观相关的幅值响应,NBI提供结构敏感的相位响应。相比传统实数或对称多模态融合,该方法显式建模了两种模态的不同角色,并抑制局部错位区域中的不可靠交互。在配对的WLI/NBI内镜数据集上的实验表明,所提可靠性感知注册引导与复数域融合显著提升病灶分割性能。模态角色互换与模块消融实验进一步验证了模态角色设计与可靠性引导跨模态交互的必要性。

原文摘要 · Abstract (English)

White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spatially misaligned due to viewpoint changes, tissue deformation, and sequential handheld acquisition. This makes direct WLI/NBI fusion prone to mixing non-corresponding regions and may even degrade segmentation around lesion boundaries. To address this problem, we propose a reliability-aware complex-domain fusion framework for paired-but-unregistered WLI/NBI lesion segmentation. The framework first establishes topology-regularized feature correspondence and further estimates where the cross-modal correspondence is reliable. Guided by this reliability, the model selectively fuses WLI and NBI features in a learnable complex representation. In this representation, WLI-derived cues mainly provide appearance-related magnitude responses, while NBI-derived cues provide structure-sensitive phase responses. Unlike conventional real-valued or symmetric multimodal fusion, the proposed method explicitly models the different roles of WLI and NBI and suppresses unreliable cross-modal interaction in locally mismatched regions. Experiments on paired WLI/NBI endoscopic datasets show that the proposed reliability-aware registration grounding and complex-domain fusion consistently improve lesion segmentation performance. Role-reversal and module ablation studies further validate the necessity of both the modality-role design and reliability-guided cross-modal interaction.

医学图像多模态融合内镜分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。