用特征函数衡量分布偏移,提升高维数据域自适应性能
CFL: On the Use of Characteristic Function Loss for Domain Alignment in Machine Learning
- 引入特征函数作为频域方法度量高维分布偏移
- 在多个数据集上显著降低域间差异,提升模型泛化能力
- 适合关注域自适应与分布对齐的研究者
机器学习模型在实际应用中常因分布偏移问题表现不佳,尤其在高风险场景下可能引发灾难性后果。以往研究多采用统计方法(如KL散度、K-S检验、Wasserstein距离)量化分布偏移。本文提出,将特征函数(Characteristic Function, CF)作为频域方法,可有效度量高维空间中的分布偏移,并用于域自适应任务。实验表明,基于特征函数的损失函数能更准确捕捉分布差异,在多个基准数据集上实现更好的域对齐效果,为解决分布偏移提供了新的有力工具。
原文摘要 · Abstract (English)
Machine Learning (ML) models are extensively used in various applications due to their significant advantages over traditional learning methods. However, the developed ML models often underperform when deployed in the real world due to the well-known distribution shift problem. This problem can lead to a catastrophic outcomes when these decision-making systems have to operate in high-risk applications. Many researchers have previously studied this problem in ML, known as distribution shift problem, using statistical techniques (such as Kullback-Leibler, Kolmogorov-Smirnov Test, Wasserstein distance, etc.) to quantify the distribution shift. In this letter, we show that using Characteristic Function (CF) as a frequency domain approach is a powerful alternative for measuring the distribution shift in high-dimensional space and for domain adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。