用代理变量校准调查数据中的系统误差,提升损失评估准确性
Proxy-Guided Measurement Calibration
- 通过因果图分离真实结果与偏差因素,利用不依赖偏差的代理变量识别误差
- 采用变分自编码器分离内容与偏差隐变量,实现对系统误差的量化估计
- 在灾损报告等真实场景中验证有效,适合做社会调查数据校正的研究者
通过调查和行政记录收集的聚合结果变量常存在系统性测量误差。例如,灾害损失数据库中,县级损失报告值可能因实地数据采集能力、报告习惯和事件特征差异而偏离真实损失。这种误校准会干扰后续分析与决策。本文研究结果误校准问题,提出一种由代理变量引导的校准框架。通过因果图将驱动真实结果的潜在内容变量与引发系统误差的潜在偏差变量分离。核心思想是:依赖真实结果但与偏差机制无关的代理变量可提供识别偏差的信息。基于此结构,提出两阶段方法,利用变分自编码器解耦内容与偏差隐变量,从而估计偏差对目标结果的影响。分析了方法假设,并在合成数据、源自随机试验的半合成数据集及真实灾损报告案例研究中进行了评估。
原文摘要 · Abstract (English)
Aggregate outcome variables collected through surveys and administrative records are often subject to systematic measurement error. For instance, in disaster loss databases, county-level losses reported may differ from the true damages due to variations in on-the-ground data collection capacity, reporting practices, and event characteristics. Such miscalibration complicates downstream analysis and decision-making. We study the problem of outcome miscalibration and propose a framework guided by proxy variables for estimating and correcting the systematic errors. We model the data-generating process using a causal graph that separates latent content variables driving the true outcome from the latent bias variables that induce systematic errors. The key insight is that proxy variables that depend on the true outcome but are independent of the bias mechanism provide identifying information for quantifying the bias. Leveraging this structure, we introduce a two-stage approach that utilizes variational autoencoders to disentangle content and bias latents, enabling us to estimate the effect of bias on the outcome of interest. We analyze the assumptions underlying our approach and evaluate it on synthetic data, semi-synthetic datasets derived from randomized trials, and a real-world case study of disaster loss reporting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。