用回归检验调查数据可信度,兼顾隐私保护与准确性。
Testing Credibility of Public and Private Surveys through the Lens of Regression
- 基于线性回归设计算法,验证样本是否代表总体。
- 在差分隐私下仍能保证回归结果可靠,误差达最优水平。
- 适合关注数据可信性与隐私保护的研究者使用。
验证样本调查是否真实反映总体,是确保后续研究有效性的关键问题。本文设计一种基于线性回归的算法,用于检验样本调查的可信度,即判断该调查能否保障回归分析结果的正确性。针对数据隐私问题,进一步考虑局部差分隐私(LDP)下的调查数据,扩展算法以适应隐私保护后的分析。所提方法还能从任意次指数分布噪声中学习回归模型,并在理论上达到ℓ₁回归的最优估计误差。算法在理论层面证明了正确性,同时降低对样本量的需求。实验在真实与合成数据上验证了其性能。
原文摘要 · Abstract (English)
Testing whether a sample survey is a credible representation of the population is an important question to ensure the validity of any downstream research. While this problem, in general, does not have an efficient solution, one might take a task-based approach and aim to understand whether a certain data analysis tool, like linear regression, would yield similar answers both on the population and the sample survey. In this paper, we design an algorithm to test the credibility of a sample survey in terms of linear regression. In other words, we design an algorithm that can certify if a sample survey is good enough to guarantee the correctness of data analysis done using linear regression tools. Nowadays, one is naturally concerned about data privacy in surveys. Thus, we further test the credibility of surveys published in a differentially private manner. Specifically, we focus on Local Differential Privacy (LDP), which is a standard technique to ensure privacy in surveys where the survey participants might not trust the aggregator. We extend our algorithm to work even when the data analysis has been done using surveys with LDP. In the process, we also propose an algorithm that learns with high probability the guarantees a linear regression model on a survey published with LDP. Our algorithm also serves as a mechanism to learn linear regression models from data corrupted with noise coming from any subexponential distribution. We prove that it achieves the optimal estimation error bound for $\ell_1$ linear regression, which might be of broader interest. We prove the theoretical correctness of our algorithms while trying to reduce the sample complexity for both public and private surveys. We also numerically demonstrate the performance of our algorithms on real and synthetic datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。