不同概率模型对低质数据的鲁棒性差异显著,自回归模型最稳定。
Robustness of Probabilistic Models to Low-Quality Data: A Multi-Perspective Analysis
- 从信息论与学习理论多角度分析模型鲁棒性机制
- GPT-2在50%标记损坏下仅测试负对数似然升至3.59
- 扩散模型图像标签一致性下降56.81%,适合关注数据质量的研究者
对现代概率模型在低质数据下的表现进行系统性对比研究,发现其鲁棒性差异显著。自回归语言模型(如GPT-2)在50%标记损坏下仍具较强韧性,测试负对数似然仅从2.87升至3.59;而同类条件下,类别条件扩散模型的图像-标签一致性相对基线下降56.81%;分类器受数据污染影响中等,且随数据集规模增大而减弱。通过信息论、PAC学习和梯度动力学多视角分析,揭示鲁棒性受两大因素主导:条件信息的丰富程度,以及训练数据的绝对信息量,前者约束学习难度,后者使有效信号压倒噪声。
原文摘要 · Abstract (English)
A systematic, comparative investigation into the effects of low-quality data reveals a stark spectrum of robustness across modern probabilistic models. We find that autoregressive language models, from token prediction to sequence-to-sequence tasks, are remarkably resilient (for GPT-2, test NLL increases modestly from 2.87 to 3.59 despite 50% token corruption). By contrast, under the same levels of data corruption, class-conditional diffusion models degrade catastrophically (image-label consistency plummets by 56.81% relative to baseline), while classifiers show a moderate impact that diminishes with dataset scale. To explain these discrepancies, we analyze the results through a multi-perspective lens, integrating information theory, PAC learning, and gradient dynamics. These analyses suggest that robustness is heavily influenced by two key principles: the richness of conditioning information, which constrains the learning problem, and the absolute information content of the training data, which allows the signal from correct information to dominate statistical noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。