在预算有限下优化多源数据采集,提升估计精度。
Learning from Biased and Costly Data Sources: Minimax-optimal Data Collection under a Budget
- 设计最优采样策略,最大化有效样本量。
- 理论证明可实现预算约束下的最小最大风险最优。
- 适用于医疗、政治调查等高成本异质数据场景。
数据收集是现代统计与机器学习流程的关键环节,尤其当需从多个异质数据源获取目标人群信息时。许多实际应用中(如医学研究或民意调查),不同数据源的采样成本各异,且观测值常带有群体标识(如健康指标、人口特征或政治立场),这些群体构成在各数据源间及与目标总体间可能存在显著差异。本文研究固定预算下的多源数据收集问题,聚焦于总体均值与组条件均值的估计。我们发现,简单匹配目标分布或使用样本均值等常规策略可能严重低效。为此,提出一种采样计划,以最大化有效样本量——即总样本量除以 $D_{χ^2}(q elmid elmidar{p}) + 1$,其中 $q$ 为目标分布,$ar{p}$ 为聚合源分布,$D_{χ^2}$ 为 $χ^2$-散度。该策略与经典后分层估计器结合,并给出了风险上界。通过提供匹配的下界,证明该方法达到预算约束下的最小最大风险最优。相关技术亦可推广至预测问题,实现对昂贵且异质数据源的合理建模。
原文摘要 · Abstract (English)
Data collection is a critical component of modern statistical and machine learning pipelines, particularly when data must be gathered from multiple heterogeneous sources to study a target population of interest. In many use cases, such as medical studies or political polling, different sources incur different sampling costs. Observations often have associated group identities - for example, health markers, demographics, or political affiliations - and the relative composition of these groups may differ substantially, both among the source populations and between sources and target population. In this work, we study multi-source data collection under a fixed budget, focusing on the estimation of population means and group-conditional means. We show that naive data collection strategies (e.g. attempting to "match" the target distribution) or relying on standard estimators (e.g. sample mean) can be highly suboptimal. Instead, we develop a sampling plan which maximizes the effective sample size - the total sample size divided by $D_{χ^2}(q\mid\mid\overline{p}) + 1$, where $q$ is the target distribution, $\overline{p}$ is the aggregated source distribution, and $D_{χ^2}$ is the $χ^2$-divergence. We pair this sampling plan with a classical post-stratification estimator and upper bound its risk. We provide matching lower bounds, establishing that our approach achieves the budgeted minimax optimal risk. Our techniques also extend to prediction problems when minimizing the excess risk, providing a principled approach to multi-source learning with costly and heterogeneous data sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。