arXiv:2510.21303cs.LG2025-10

数据差异如何影响模型多样性,这篇论文给出了新视角。

Data as a Lever: A Neighbouring Datasets Perspective on Predictive Multiplicity

  • 从邻近数据集角度重构数据处理,揭示数据分布对模型多样性的影响。
  • 发现类别间分布重叠越高的邻近数据集,模型多样性越低。
  • 提出可感知多样性的主动学习与数据补全方法,适合关注模型稳定性的研究者。

多重性(Multiplicity)指存在多个表现相当但相互竞争的模型,近年来受到越来越多关注。以往研究多聚焦于模型选择,却忽视了数据在塑造多重性中的关键作用。本文提出邻近数据集框架,认为许多数据处理过程本质上是在不同邻近数据集间做选择。在此框架下,我们发现一个反直觉的理论关系:类别间分布重叠越高的邻近数据集,其模型多重性越低。基于此,我们将该框架拓展至主动学习和数据补全两个领域,首次系统研究现有算法中的多重性问题,并提出两种新型多重性感知方法:多重性感知的主动学习数据采集策略与多重性感知的数据补全方法。

原文摘要 · Abstract (English)

Multiplicity, the existence of equally good yet competing models, has received growing attention in recent years. While prior work has emphasized modelling choices, the critical role of data in shaping multiplicity has been largely overlooked. In this work, we first introduce a neighbouring datasets framework, arguing that much of data processing can be reframed as choosing between neighbouring datasets. Under this framework, we find a counterintuitive theoretical relationship: neighbouring datasets with greater inter-class distribution overlap exhibit lower multiplicity. Building on this insight, we apply our framework to two domains: active learning and data imputation. For each, we establish natural extensions of the neighbouring datasets perspective, conduct the first systematic study of multiplicity in existing algorithms, and finally, propose novel multiplicity-aware methods, namely, multiplicity-aware data acquisition strategies for active learning and multiplicity-aware data imputation.

多重性数据处理主动学习数据补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。