arXiv:2605.05084cs.LG2026-05被引 3

通过重排数据顺序降低域适应中的误差,提升模型泛化能力。

Order Matters: Improving Domain Adaptation by Reordering Data

  • 重排训练数据采样顺序以减少域差异估计的方差
  • 在两个图像分类基准上实现更高的目标域准确率
  • 适合关注域适应稳定性和性能提升的研究者

域偏移仍是机器学习模型部署到真实世界的主要挑战。无监督域适应(UDA)旨在通过训练过程中最小化域差异来应对这一问题,但其差异估计在随机设置下存在高方差,可能削弱方法的理论优势。本文提出一种新的无偏随机方差减少技术——ORDERED,通过优化训练数据的采样顺序来降低差异估计误差。针对两种特定的域差异损失(相关对齐与最大均值差异),将它们的随机估计误差建模为数据采样顺序的函数,并提出一种实用优化算法。仿真结果表明,相比现有方法方差显著降低;在两个域偏移图像分类基准上的实验也验证了目标域准确率的提升。

原文摘要 · Abstract (English)

Domain shift remains a key challenge in deploying machine learning models to the real world. Unsupervised domain adaptation (UDA) aims to address this by minimising domain discrepancy during training, but the discrepancy estimates suffer from high variance in stochastic settings, which can stifle the theoretical benefits of the method. This paper proposes Optimal Reordering of Data for Error-Reduced Estimation of Discrepancy (ORDERED), a novel unbiased stochastic variance reduction technique which reduces the discrepancy estimation error by optimising the order in which the training data are sampled. We consider two specific domain discrepancy losses (correlation alignment and the maximum mean discrepancy), formulate their stochastic estimation error as a function of the data sampling order, and propose a practical optimisation algorithm. Our simulations demonstrate reduced variance compared to related methods, and experiments on two domain shift image classification benchmarks show improved target domain accuracy.

域适应数据重排方差减少

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。