arXiv:2508.20616cs.LGstat.ML2025-08综述

不依赖维度的调查数据可信度检测方法,提升高维数据评估效率

Dimension Agnostic Testing of Survey Data Credibility through the Lens of Regression

  • 基于回归模型设计特定距离度量,聚焦可信度验证而非建模
  • 算法样本复杂度与数据维度无关,理论证明高效可靠
  • 适合高维调查数据分析,尤其适用于大规模社会研究

评估样本调查是否可信地代表总体,是保障后续研究有效性的关键。传统方法需估计高维分布间距离,样本需求随维度指数增长。但某些分析模型下,结论在不同分布间仍保持一致。本文提出一种任务导向的可信度评估方法,针对回归模型定义特定距离度量,并设计算法验证调查数据可信性。该算法样本复杂度与数据维度无关,因其仅关注可信度验证而非重建回归模型。反观重建模型的方法,其样本复杂度随维度线性增长。我们证明了算法的理论正确性,并通过数值实验验证了其性能。

原文摘要 · Abstract (English)

Assessing whether a sample survey credibly represents the population is a critical question for ensuring the validity of downstream research. Generally, this problem reduces to estimating the distance between two high-dimensional distributions, which typically requires a number of samples that grows exponentially with the dimension. However, depending on the model used for data analysis, the conclusions drawn from the data may remain consistent across different underlying distributions. In this context, we propose a task-based approach to assess the credibility of sampled surveys. Specifically, we introduce a model-specific distance metric to quantify this notion of credibility. We also design an algorithm to verify the credibility of survey data in the context of regression models. Notably, the sample complexity of our algorithm is independent of the data dimension. This efficiency stems from the fact that the algorithm focuses on verifying the credibility of the survey data rather than reconstructing the underlying regression model. Furthermore, we show that if one attempts to verify credibility by reconstructing the regression model, the sample complexity scales linearly with the dimensionality of the data. We prove the theoretical correctness of our algorithm and numerically demonstrate our algorithm's performance.

调查可信度高维数据回归模型样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。