对比多种数据诊断信号,发现现有方法难检测异常样本。
A Comparative Analysis of Influence Signals for Data Debugging
- 基于影响度信号比较不同数据噪声类型检测能力
- 自影响信号可有效识别标签错误样本,但无法检测异常
- 揭示了影响度衰减与符号抵消问题,适合模型调试研究者
提升训练样本质量对增强机器学习模型的可靠性与性能至关重要。本文对基于影响度的数据诊断信号进行了对比评估,这些信号可在构建模型过程中识别标签错误和异常样本,从而减少对专用缺陷检测器的依赖。尽管近年来已有多种影响度信号(如自影响、平均绝对影响、边际影响、GD-class)被提出,但尚无在统一影响估计器(如TraceIn)下,针对图像与表格数据模态及从零训练或基础模型训练的深度学习模型,系统评估其对不同类型数据缺陷(如标签错误与异常)检测能力的研究。通过大量实验,我们发现自影响等信号能有效检测标签错误样本,但现有信号均无法检测异常样本。此外,现有信号未考虑训练动态过程中的影响变化,部分信号还存在影响度抵消效应——因无符号分数累加导致得分归零,造成误导性影响归属。
原文摘要 · Abstract (English)
Improving the quality of training samples is crucial for improving the reliability and performance of ML models. In this paper, we conduct a comparative evaluation of influence-based signals for debugging training data. These signals can potentially identify both mislabeled and anomalous samples from a potentially noisy training set as we build the models and hence alleviate the need for dedicated glitch detectors. Although several influence-based signals (e.g., Self-Influence, Average Absolute Influence, Marginal Influence, GD-class) have been recently proposed in the literature, there are no experimental studies for assessing their power in detecting different glitch types (e.g., mislabeled and anomalous samples) under a common influence estimator (e.g., TraceIn) for different data modalities (image and tabular), and deep learning models (trained from scratch or foundation). Through extensive experiments, we show that signals like Self-Influence effectively detect mislabeled samples, but none of the existing signals can detect anomalies. Existing signals do not take into account the training dynamics, i.e., how the samples' influence on the model changes during training, while some signals fall into influence cancellation effects, i.e., influence score is zero due to unsigned scores accumulation, resulting in misleading influence attribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。