arXiv:2512.10952cs.LGcs.AI2025-12AAAI

从多个数据源中筛选高质量数据集,提升模型性能。

Hierarchical Dataset Selection for High-Quality Data Sharing

  • 按数据集和分组层级建模数据价值,实现高效选择。
  • 在两个基准上最高提升26.2%准确率,探索步骤更少。
  • 适合资源受限或数据源多的现实场景,可扩展性强。

现代机器学习的成功依赖于高质量训练数据。在实际场景中,如从公开仓库获取数据或跨机构共享,数据天然以独立数据集形式存在,且各数据集在相关性、质量与实用性上差异显著。现有方法多关注个体样本选择,忽略数据集及来源间的差异。本文正式定义数据集选择任务:在资源受限下,从异构数据池中选择整套数据集以提升下游性能。提出DaSH方法,通过在数据集与分组(如机构、集合)层面建模效用,实现有限观测下的高效泛化。在Digit-Five和DomainNet两个公开基准上,DaSH相比最优基线最高提升26.2%准确率,且所需探索步骤显著减少。消融实验表明,该方法对低资源环境和缺乏相关数据集情况均具鲁棒性,适用于实际多源学习流程中的可扩展、自适应数据集选择。

原文摘要 · Abstract (English)

The success of modern machine learning hinges on access to high-quality training data. In many real-world scenarios, such as acquiring data from public repositories or sharing across institutions, data is naturally organized into discrete datasets that vary in relevance, quality, and utility. Selecting which repositories or institutions to search for useful datasets, and which datasets to incorporate into model training are therefore critical decisions, yet most existing methods select individual samples and treat all data as equally relevant, ignoring differences between datasets and their sources. In this work, we formalize the task of dataset selection: selecting entire datasets from a large, heterogeneous pool to improve downstream performance under resource constraints. We propose Dataset Selection via Hierarchies (DaSH), a dataset selection method that models utility at both dataset and group (e.g., collections, institutions) levels, enabling efficient generalization from limited observations. Across two public benchmarks (Digit-Five and DomainNet), DaSH outperforms state-of-the-art data selection baselines by up to 26.2% in accuracy, while requiring significantly fewer exploration steps. Ablations show DaSH is robust to low-resource settings and lack of relevant datasets, making it suitable for scalable and adaptive dataset selection in practical multi-source learning workflows.

数据选择多源学习高效采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。