通过分层采样降低域适应中的方差,提升模型泛化能力。
Variance Matters: Improving Domain Adaptation via Stratified Sampling
- 采用分层采样策略减少域间差异估计的方差
- 在四个数据集上显著提升域适应性能和差异估计精度
- 适合关注模型鲁棒性与真实场景部署的研究者
领域偏移仍是机器学习模型落地现实世界的主要挑战。无监督域适应(UDA)旨在通过训练过程中最小化领域差异来应对这一问题,但其差异估计在随机环境下方差过高,削弱了方法的理论优势。本文提出首个专用于UDA的随机方差缩减技术——分层采样域适应(VaRDASS)。针对相关性对齐与最大均值差异(MMD)两种具体度量,推导出相应的分层目标函数。我们给出了期望误差与最坏情况下的误差界,并证明在特定假设下,所提MMD目标函数理论上最优(即最小化方差)。最后,设计了一种类k-means的实用优化算法并进行分析。在四个领域偏移数据集上的实验表明,该方法显著提升了差异估计准确性和目标域性能。
原文摘要 · Abstract (English)
Domain shift remains a key challenge in deploying machine learning models to the real world. Unsupervised domain adaptation (UDA) aims to address this by minimising domain discrepancy during training, but the discrepancy estimates suffer from high variance in stochastic settings, which can stifle the theoretical benefits of the method. This paper proposes Variance-Reduced Domain Adaptation via Stratified Sampling (VaRDASS), the first specialised stochastic variance reduction technique for UDA. We consider two specific discrepancy measures -- correlation alignment and the maximum mean discrepancy (MMD) -- and derive ad hoc stratification objectives for these terms. We then present expected and worst-case error bounds, and prove that our proposed objective for the MMD is theoretically optimal (i.e., minimises the variance) under certain assumptions. Finally, a practical k-means style optimisation algorithm is introduced and analysed. Experiments on four domain shift datasets demonstrate improved discrepancy estimation accuracy and target domain performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。