arXiv:2605.16081cs.LGcs.CV2026-05中稿 · ICML

解决预训练模型带来的系统性标注噪声问题

MIND: Decoupling Model-Induced Label Noise via Latent Manifold Disentanglement

论文配图:MIND: Decoupling Model-Induced Label Noise via Latent Manifold Disentanglement
图 1 · 摘自论文原文
  • 通过潜在流形解耦分离噪声与特征结构
  • 在CIFAR-100和3D数据集上显著提升鲁棒性
  • 适合用于大模型蒸馏与真实场景数据纠错

基于预训练模型的自动标注范式在数据密集型应用中占主导地位,但引入了关键挑战:模型诱导的标注噪声。这种噪声源于标注器的归纳偏置,表现为与局部特征流形紧密耦合的系统性错误。现有方法依赖全局转移矩阵,难以捕捉此类结构性模式;而学习实例级矩阵则数学上不可行。本文提出模型诱导噪声解耦(MIND)框架,理论上证明高维噪声流形可通过潜在流形解耦分解为可处理的子空间分量。具体地,潜在线性解耦估计器(LDE)动态将样本投影至具有一致错误模式的潜在结构聚类中,实现无需真值锚点的噪声可识别性。为严格评估鲁棒性,采用分层协议:从CIFAR-100上的可控噪声测试,到大规模真实3D数据集(S3DIS、ScanNet)的结构压力测试,其中错误模式明确耦合于几何流形。实验表明,MIND在复杂基准上显著优于当前最优方法,并有效纠正视觉语言模型(如OpenSeg)的零样本幻觉,展现出作为大模型鲁棒蒸馏框架的巨大潜力。

原文摘要 · Abstract (English)

The paradigm of learning from automatic annotations driven by pre-trained experts and Foundation Models dominates data-hungry applications. However, it introduces a critical challenge: model-induced label noise. Unlike stochastic noise in classical robust learning, this noise stems from annotator inductive biases, manifesting as systematic errors tightly coupled with local feature manifolds. Existing methods relying on global transition matrices underfit these structural patterns, while learning instance-specific matrices remains mathematically intractable. We propose Model-Induced Noise Decoupling (MIND), a theoretically grounded framework addressing this dilemma. We demonstrate that the high-dimensional noise manifold can be decoupled into tractable, subspace-dependent components via Latent Manifold Disentanglement. Specifically, our Latent Decoupling Estimator (LDE) dynamically projects samples into latent structural clusters with consistent error modes, facilitating noise identifiability without ground-truth anchor points. To rigorously evaluate robustness, we adopt a hierarchical protocol: moving from controlled noise on CIFAR-100 to a structural stress test on large-scale real-world 3D datasets (S3DIS, ScanNet), where error patterns explicitly couple with geometric manifolds. Empirically, MIND significantly outperforms state-of-the-art methods on these complex benchmarks and effectively corrects zero-shot hallucinations from Vision-Language Models (e.g., OpenSeg), highlighting its potential as a robust distillation framework for Foundation Models.

噪声去除大模型蒸馏3D理解流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。