arXiv:2506.20573stat.MLcs.LG2025-06

提出无需针对特定模型的鲁棒数据过滤方法,提升多模型通用性。

LARP: Learner-Agnostic Robust Data Prefiltering

  • 设计不依赖下游模型的统一数据过滤策略
  • 理论证明最差情况损失有上界,但性能略低于定制化过滤
  • 适合需批量处理数据、减少重复清洗的团队使用

公开数据集对现代机器学习和统计推断至关重要,但常包含低质量或污染样本,可能损害模型性能。为此,亟需一种可由数据提供方应用的合理预过滤方法,以同时保护多种下游统计与学习方法的准确性。本文形式化并分析了「学习者无关鲁棒数据预过滤」(LARP)问题,即在预设学习者集合上,设计具有最坏情况损失保证的预过滤程序。我们在两个理论设置中证明了 LARP 的可行性,给出了最坏情况损失的上界。理论结果表明,通过 LARP 保护异构学习者集合会带来一定性能损失,我们称其为「LARP 代价」。通过图像与表格任务的实验,我们实证测量了该代价。此外,在一个博弈论模型中,我们探讨了 LARP 在节省重复数据整理成本方面的潜在优势,其中下游学习者可分摊单一预过滤的成本。

原文摘要 · Abstract (English)

Public datasets, crucial for modern machine learning and statistical inference, often contain low-quality or contaminated samples that can harm model performance. This creates a need for principled prefiltering procedures that a data provider can apply to protect the accuracy of a range of potential downstream statistical and learning procedures simultaneously. In this work, we formalize and analyze Learner-Agnostic Robust data Prefiltering (LARP), the problem of designing prefiltering procedures with guarantees on the worst-case loss over a pre-specified set of learners. We establish the feasibility of LARP in two theoretical settings, by providing upper-bound guarantees on the worst-case loss. Our theoretical results indicate that protecting heterogeneous learner sets via LARP comes at the price of some performance loss compared to individual, learner-specific prefiltering; we call this gap the price of LARP. To assess this gap in performance, we empirically measure the price of LARP across image and tabular tasks. We further explore potential benefits of LARP from the perspective of saving on repeated data curation efforts, in a game-theoretic model where the downstream learners can split the cost of the single prefiltering.

数据清洗鲁棒学习通用过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。