arXiv:2602.16125cs.LG2026-02

筛选优质数据源可显著提升共享特征提取效果

On the Power of Source Screening for Learning Shared Feature Extractors

  • 通过筛选高信息量数据子集,实现更优的共享特征学习
  • 仅用部分数据即可达到理论最优性能,丢弃率可高达60%
  • 适合需要高效利用异构数据的研究者参考

共享表示学习被广泛认为是分离多源异质数据共性与差异的有效方法。现有工作通常同时训练通用特征提取器和源特定头模块,但低相关性或低质量数据源可能损害表示学习效果。本文聚焦于传统上被视为‘优质’的数据源集合——各源对真实共性结构具有相似相关性和质量。在可处理的线性设定下,即各源共享低维子空间,我们发现源筛选在统计最优子空间估计中起核心作用。对于一大类问题实例,精心挑选的部分源联合训练即可实现极小极大最优,即使舍弃大量数据也无影响。论文定义了‘信息丰富子群体’概念,提出识别算法与实用启发式方法,并通过合成及真实数据集的理论分析与实证评估验证其有效性。

原文摘要 · Abstract (English)

Learning with shared representation is widely recognized as an effective way to separate commonalities from heterogeneity across various heterogeneous sources. Most existing work includes all related data sources via simultaneously training a common feature extractor and source-specific heads. It is well understood that data sources with low relevance or poor quality may hinder representation learning. In this paper, we further dive into the question of which data sources should be learned jointly by focusing on the traditionally deemed ``good'' collection of sources, in which individual sources have similar relevance and qualities with respect to the true underlying common structure. Towards tractability, we focus on the linear setting where sources share a low-dimensional subspace. We find that source screening can play a central role in statistically optimal subspace estimation. We show that, for a broad class of problem instances, training on a carefully selected subset of sources suffices to achieve minimax optimality, even when a substantial portion of data is discarded. We formalize the notion of an informative subpopulation, develop algorithms and practical heuristics for identifying such subsets, and validate their effectiveness through both theoretical analysis and empirical evaluations on synthetic and real-world datasets.

特征提取数据筛选共享表示子空间学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。