系统梳理分布外数据的应对方法,揭示现有模型在分布漂移下的失效原因。
Handling Out-of-Distribution Data: A Survey
- 区分特征漂移与概念漂移,提出统一应对框架
- 总结检测、度量与缓解分布漂移的主流技术
- 适合关注模型鲁棒性与真实场景泛化的研究者
机器学习与数据驱动应用中,训练与部署阶段的数据分布变化(即分布漂移)是重大挑战。本文系统分析两类主要分布漂移:(i) 特征漂移(协变量漂移),即训练与测试数据中特征值分布改变;(ii) 概念/语义漂移,即模型在测试阶段遭遇训练中未见的新类别,导致学习概念发生变化。本文贡献有三:首先,形式化分布漂移问题,指出传统方法难以有效应对,呼吁发展能同时适应各类漂移的模型;其次,阐明处理分布漂移的重要性,全面综述检测、度量与缓解漂移的技术进展;最后,评估当前应对机制现状,并提出未来研究方向。整体上,本文回顾了分布漂移相关文献,特别聚焦于以往综述中被忽视的分布外数据问题。
原文摘要 · Abstract (English)
In the field of Machine Learning (ML) and data-driven applications, one of the significant challenge is the change in data distribution between the training and deployment stages, commonly known as distribution shift. This paper outlines different mechanisms for handling two main types of distribution shifts: (i) Covariate shift: where the value of features or covariates change between train and test data, and (ii) Concept/Semantic-shift: where model experiences shift in the concept learned during training due to emergence of novel classes in the test phase. We sum up our contributions in three folds. First, we formalize distribution shifts, recite on how the conventional method fails to handle them adequately and urge for a model that can simultaneously perform better in all types of distribution shifts. Second, we discuss why handling distribution shifts is important and provide an extensive review of the methods and techniques that have been developed to detect, measure, and mitigate the effects of these shifts. Third, we discuss the current state of distribution shift handling mechanisms and propose future research directions in this area. Overall, we provide a retrospective synopsis of the literature in the distribution shift, focusing on OOD data that had been overlooked in the existing surveys.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。