arXiv:2511.00997cs.CV2025-11

无需配对数据,用迭代去噪法自动清除复杂噪声。

MID: A Self-supervised Multimodal Iterative Denoising Framework

  • 通过迭代加噪学习噪声估计与去除网络。
  • 在4个视觉任务中达到最先进性能。
  • 适用于生物医学等非图像领域,无需干净数据标注。

数据去噪在科学与工程领域长期面临挑战。真实世界数据常受复杂非线性噪声污染,传统基于规则的去噪方法难以应对。为此,我们提出一种新型自监督多模态迭代去噪(MID)框架。MID将采集到的噪声数据建模为非线性噪声累积过程中的一个状态,通过迭代引入更多噪声,学习两个神经网络:一个用于估计当前噪声步数,另一个用于预测并减去相应的噪声增量。针对复杂非线性污染,MID采用一阶泰勒展开对噪声过程局部线性化,实现有效迭代去除。关键优势在于无需成对的清洁-噪声数据集,可直接从噪声输入中学习噪声特性。在四个经典计算机视觉任务上的实验表明,MID具备强鲁棒性、高适应性,并持续取得最先进性能;此外,在生物医学与生物信息学任务中也展现出优异表现。

原文摘要 · Abstract (English)

Data denoising is a persistent challenge across scientific and engineering domains. Real-world data is frequently corrupted by complex, non-linear noise, rendering traditional rule-based denoising methods inadequate. To overcome these obstacles, we propose a novel self-supervised multimodal iterative denoising (MID) framework. MID models the collected noisy data as a state within a continuous process of non-linear noise accumulation. By iteratively introducing further noise, MID learns two neural networks: one to estimate the current noise step and another to predict and subtract the corresponding noise increment. For complex non-linear contamination, MID employs a first-order Taylor expansion to locally linearize the noise process, enabling effective iterative removal. Crucially, MID does not require paired clean-noisy datasets, as it learns noise characteristics directly from the noisy inputs. Experiments across four classic computer vision tasks demonstrate MID's robustness, adaptability, and consistent state-of-the-art performance. Moreover, MID exhibits strong performance and adaptability in tasks within the biomedical and bioinformatics domains.

去噪自监督多模态迭代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。