arXiv:2605.01874cs.LGcs.AI2026-05

利用数据对称性提升噪声数据中优质子集的筛选效果

Leveraging Data Symmetries to Select an Optimal Subset of Training Data under Label Noise

论文配图:Leveraging Data Symmetries to Select an Optimal Subset of Training Data under Label Noise
图 1 · 摘自论文原文
  • 基于数据对称性改进k-NN,增强噪声环境下样本筛选能力
  • 在高维场景下使分类器性能逼近贝叶斯最优水平
  • 即使对称性信息不完整,学习到的不变表示仍能有效选优

机器学习模型性能依赖大规模标注数据,但来自不同来源的数据常含标签噪声。已有研究表明,在噪声环境下,可能存在一个子集,使模型性能接近在无噪声数据上训练的效果。目前常用方法cutstats通过k近邻(k-NN)识别低噪声样本,但在高维数据上的表现尚不明确。本文首次从理论上证明,使用cutstats选取子集的分类器性能受k-NN准确率影响。进一步表明,在噪声环境中,利用数据不变性与底层对称性可显著提升k-NN性能,使其在高维情形下趋近贝叶斯最优分类器。最后,针对现实场景中对称性信息部分已知的情况,我们证明学习到的不变表示仍能有效辅助识别近似最优训练子集。

原文摘要 · Abstract (English)

The performance of machine learning models often relies on large labeled datasets; however, data collected from diverse sources can contain label noise. Recent work has shown that, in noisy settings, there may exist a subset of the training data on which models can achieve performance comparable to training on a noise-free dataset. A widely used method for identifying such subsets is cutstats, which employs k-nearest neighbors (k-NN) to detect low-noise samples. However, its performance on high-dimensional data remains largely unexplored. In this work, we formally establish that the performance of a classifier trained on a subset of a noisy dataset selected via cutstats is influenced by the accuracy of k-NN. We further demonstrate that, in noisy environments, exploiting data invariance and knowledge of underlying symmetries can significantly enhance the performance of k-NN, bringing it closer to the Bayes optimal classifier even in high-dimensional regimes. Finally, we show that for real-world scenarios, where information about the underlying invariance is only partially known, learnt invariant representations can still facilitate the identification of near-optimal subsets.

数据清洗标签噪声不变性k-NN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。