arXiv:2505.16159cs.LG2025-05被引 1

模型为何能在错误标注下仍保持准确?

Why Can Accurate Models Be Learned from Inaccurate Annotations?

  • 分析权重矩阵的主子空间,发现噪声主要影响低奇异分量
  • 在适度错误率下,主子空间与干净数据训练结果高度一致
  • 提出轻量级插件LIP,提升模型对噪声标签的鲁棒性

由于精确标注成本高昂,从不准确标注中学习受到广泛关注。尽管存在错误标签,基于噪声数据训练的模型仍能做出准确预测。本文从经验与理论角度分析权重矩阵,发现标签不准确主要导致低奇异分量引入噪声,而主子空间仅受轻微扰动。在一定范围内,错误标签训练的主子空间与清洁标签训练结果保持高度对齐,保留了任务相关的核心信息。我们形式化证明了主子空间夹角在中等标签错误率下偏差极小,解释了模型的有效泛化能力。基于此,提出轻量级插件LIP,帮助分类器保留主子空间信息并抑制标签错误带来的噪声。大量实验表明,LIP在多种不准确标注场景下均能持续提升现有算法性能。研究成果为理解模型在不准确监督下的鲁棒性提供了理论与实践启示。

原文摘要 · Abstract (English)

Learning from inaccurate annotations has gained significant attention due to the high cost of precise labeling. However, despite the presence of erroneous labels, models trained on noisy data often retain the ability to make accurate predictions. This intriguing phenomenon raises a fundamental yet largely unexplored question: why models can still extract correct label information from inaccurate annotations remains unexplored. In this paper, we conduct a comprehensive investigation into this issue. By analyzing weight matrices from both empirical and theoretical perspectives, we find that label inaccuracy primarily accumulates noise in lower singular components and subtly perturbs the principal subspace. Within a certain range, the principal subspaces of weights trained on inaccurate labels remain largely aligned with those learned from clean labels, preserving essential task-relevant information. We formally prove that the angles of principal subspaces exhibit minimal deviation under moderate label inaccuracy, explaining why models can still generalize effectively. Building on these insights, we propose LIP, a lightweight plug-in designed to help classifiers retain principal subspace information while mitigating noise induced by label inaccuracy. Extensive experiments on tasks with various inaccuracy conditions demonstrate that LIP consistently enhances the performance of existing algorithms. We hope our findings can offer valuable theoretical and practical insights to understand of model robustness under inaccurate supervision.

噪声标签主子空间鲁棒学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。