研究多个分布偏移同时发生时模型的鲁棒性,发现增强数据效果最好。
An Analysis of Model Robustness across Concurrent Distribution Shifts
- 测试26种算法在168组数据上的表现,涵盖多种并发分布偏移
- 并发偏移通常比单一偏移更差,但部分模型能跨偏移泛化
- 启发式数据增强在合成与真实数据上均表现最佳
机器学习模型在源数据上优化良好,但在面对分布偏移(DS)时往往无法预测目标数据。以往基准研究多关注简单偏移,而现实中分布偏移常以复杂形式并发出现。本研究扩展分析范围,涵盖未见领域偏移与虚假相关性等多重并发偏移。在八个数据集的168个源-目标对上,评估了从简单数据增强到基础模型零样本推理共26种算法。对超过10万次模型的分析显示:(i) 并发分布偏移通常比单一偏移导致性能下降,有少数例外;(ii) 在某一偏移上提升泛化的模型,往往对其他偏移也有效;(iii) 启发式数据增强在合成与真实数据上均取得最佳整体表现。
原文摘要 · Abstract (English)
Machine learning models, meticulously optimized for source data, often fail to predict target data when faced with distribution shifts (DSs). Previous benchmarking studies, though extensive, have mainly focused on simple DSs. Recognizing that DSs often occur in more complex forms in real-world scenarios, we broadened our study to include multiple concurrent shifts, such as unseen domain shifts combined with spurious correlations. We evaluated 26 algorithms that range from simple heuristic augmentations to zero-shot inference using foundation models, across 168 source-target pairs from eight datasets. Our analysis of over 100K models reveals that (i) concurrent DSs typically worsen performance compared to a single shift, with certain exceptions, (ii) if a model improves generalization for one distribution shift, it tends to be effective for others, and (iii) heuristic data augmentations achieve the best overall performance on both synthetic and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。