arXiv:2508.11144cs.LG2025-08被引 2

针对多个小规模数据源,实现精准预测与源间差异保留的新型迁移学习方法。

CTRL Your Shift: Clustered Transfer Residual Learning for Many Small Datasets

  • 通过聚类与残差迁移结合,自适应处理多源分布偏移问题。
  • 在5个大规模数据集上优于现有基准,尤其在小样本源上提升显著。
  • 适合需要差异化预测的场景,如难民安置、医疗分组等政策决策。

机器学习任务常涉及来自多个不同来源(如地理位置、治疗组别)的大规模数据。在此类场景中,不仅要求整体预测准确,还需确保各来源内部预测可靠,并保留跨来源间关键差异。例如,瑞士全国庇护者安置项目正试点使用基于机器学习的就业预测系统,为新到家庭在接收国内进行算法化地理分配,这需要对众多且通常样本量较小的地区生成具有信息量且区分度高的预测结果。然而,此类任务面临多重挑战:存在大量不同数据源、源间分布偏移显著,以及各源样本量差异巨大。本文提出聚类转移残差学习(CTRL),一种元学习方法,融合跨域残差学习与自适应聚类/池化的优势,同时提升整体准确率并保持源级异质性。我们建立了新理论,证明高质量聚类可高效学习,无需反复对候选子集重训练模型。在5个大规模数据集上评估了CTRL,包括瑞士国家庇护项目的真实数据集,该系统当前正处于试点阶段。结果显示,无论使用何种基础学习器,CTRL在多个关键指标上均持续优于现有先进基准。

原文摘要 · Abstract (English)

Machine learning (ML) tasks often utilize large-scale data that is drawn from several distinct sources, such as different locations, treatment arms, or groups. In such settings, practitioners often desire predictions that not only exhibit good overall accuracy, but also remain reliable within each source and preserve the differences that matter across sources. For instance, several asylum and refugee resettlement programs now use ML-based employment predictions to guide where newly arriving families are placed within a host country, which requires generating informative and differentiated predictions for many and often small source locations. However, this task is made challenging by several common characteristics of the data in these settings: the presence of numerous distinct data sources, distributional shifts between them, and substantial variation in sample sizes across sources. This paper introduces Clustered Transfer Residual Learning (CTRL), a meta-learning method that combines the strengths of cross-domain residual learning and adaptive pooling/clustering in order to simultaneously improve overall accuracy and preserve source-level heterogeneity. We establish new theory showing that high-quality clusters can be learned efficiently, bypassing the need for repeated model refitting over candidate subsets. We evaluate CTRL alongside other state-of-the-art benchmarks on 5 large-scale datasets. This includes a dataset from the national asylum program in Switzerland, where the algorithmic geographic assignment of asylum seekers is currently being piloted. CTRL consistently outperforms the benchmarks across several key metrics and when using a range of different base learners.

迁移学习小样本聚类政策推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。