用大模型提升灾备调查质量,实测可降偏倚、补缺失数据。
Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses

- 构建基于保护动机理论的联合知识图谱,指导大模型生成合理问卷与填补数据。
- 新模型在灾后缺失数据下误差更低,偏差接近零,优于传统方法。
- 发现整体低偏差可能掩盖子群体错误,建议分组审计报告标准。
调查研究面临响应率下降、样本偏差、高危人群数据缺失及线上欺诈等结构性挑战。本文以佛罗里达州2024年飓风米尔顿灾备调查(n=946)为实证场景,提出并评估五阶段大语言模型(LLM)集成框架:问卷设计、抽样、预测试、缺失值填补与后收集分析。引入保护动机理论(PMT)约束的共现知识图谱,开发七种LLM配置,包括零样本、检索增强基线及新型理论引导变体。所提锚定边际理论引导模型(A-TLM)在灾备相关块状非随机缺失(MNAR)条件下,均方根误差(RMSE)为1.439,优于三种经典填补方法(最低1.496);且偏差仅-0.121,远低于随机森林法的-0.631。基于因果结构的检索组织优于无序检索与分步推理(MAE 0.993 vs. 1.097)。研究指出,近零总体偏差可能掩盖子群体系统性误差,建议采用子群分层偏差审计作为报告规范。检索约束的知识图谱聊天机器人表明,通过强制拒绝可有效控制幻觉。
原文摘要 · Abstract (English)
Survey research faces mounting structural challenges: declining response rates, sample bias, block-wise missingness among at-risk respondents, and AI-assisted fraudulent completions in online panels. Large language models (LLMs) have been proposed as a remedy, yet rigorous evaluations across the full survey workflow remain scarce, particularly in disaster contexts where data quality matters most. We present and evaluate a five-stage framework for LLM integration covering questionnaire design, sample selection, pilot testing, missing-data imputation, and post-collection analysis, using the 2024 Hurricane Milton preparedness survey of Florida residents (n=946) as a shared empirical testbed. We introduce a Protection Motivation Theory (PMT)-constrained co-occurrence knowledge graph and develop seven LLM configurations spanning zero-shot inference, retrieval-augmented baselines, and novel theory-informed variants. Our proposed Anchored Marginal Theory-Informed LLM (A-TLM) outperforms all three classical imputation baselines (IPW/MI, MICE+PMM, missForest) on RMSE under disaster-relevant block-wise MNAR conditions (S4 RMSE 1.439 vs. 1.496 for the next-best), while achieving near-zero signed bias (-0.121) where the random-forest imputer produces the largest absolute bias (-0.631). Organizing retrieval around PMT causal structure and integrating all evidence in a single model call outperforms unstructured retrieval and staged sequential inference (MAE 0.993 vs. 1.097 for standard RAG). We document that near-zero aggregate bias can mask opposing subgroup errors and propose subgroup-stratified bias auditing as a reporting standard. A retrieval-constrained knowledge-graph chatbot demonstrates that hallucination is architecturally manageable through grounded refusal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。