arXiv:2602.16061stat.MLcs.LG2026-02被引 1

用AI生成的辅助变量解决缺失数据识别难题,不依赖强假设。

AI-Generated Measurements for Identification and Inference with Missing Data: A Weak Shadow Variable Approach

  • 用大模型生成弱影子变量,作为缺失结果的间接信息源。
  • 在真实客服对话数据上,置信区间窄89%,误差低41%。
  • 适合缺乏完整数据、但有丰富交互记录的研究者使用。

在商业与社会科学中,结果缺失常与未观测结果相关,即缺失不是随机的(MNAR),导致总体量难以识别,除非做出强假设。如今,客户交互历史等非结构化数据日益丰富,可通过大语言模型(LLMs)生成结构化测量。本文提出一种假设轻量的局部识别框架,利用这些测量作为弱影子变量——即在真实结果和可观测协变量条件下,与缺失性条件独立且对结果具信息性的代理变量。该方法无需准确预测缺失结果,也不需满足经典影子变量中的完备性要求。我们通过一对线性规划刻画了总体量的紧界;并提出一种局部惩罚估计器,可应对抽样误差,结合子采样算法构建置信区间。在基于真实客服对话的半合成实验中,使用弱影子变量的置信区间比无辅助信息时窄89%,其中点估计误差较传统MNAR方法降低约41%。

原文摘要 · Abstract (English)

Across business and social science applications, outcomes are often missing in ways that depend on the unobserved outcomes themselves. In service systems, for example, whether a customer submits a rating depends on the rating they would have provided. Such missing-not-at-random (MNAR) mechanisms make population quantities difficult to identify without strong assumptions on the observation process. Meanwhile, rich unstructured data, such as customer interaction histories, are increasingly available and can be used to construct structured measurements using tools such as large language models (LLMs). In this work, we develop an assumption-lean partial identification framework that uses such measurements as weak shadow variables, defined as outcome-informative proxies that are conditionally independent of missingness given the true outcome and observed covariates. Importantly, they need not accurately predict missing outcomes or satisfy the completeness requirement in the classical shadow variable literature. For identification, we characterize sharp bounds on population quantities through a pair of linear programs. For estimation and inference, we propose a localized penalized estimator that remains feasible under sampling error, and a subsampling algorithm for constructing confidence intervals. In semi-synthetic experiments using real customer-service dialogues, weak-shadow-variable intervals are about 89\% narrower than those without auxiliary information, while their midpoints have around 41\% lower estimation error than classical MNAR methods.

缺失数据弱影子变量大模型应用因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。