用少量可用数据微调大模型,提升问卷缺失数据的预测准确率。
Learning from Convenience Samples: A Case Study on Fine-Tuning LLMs for Survey Non-response in the German Longitudinal Election Study
- 基于部分问卷数据微调开源大模型,模拟缺失回答
- 在随机缺失时性能媲美传统分类器,在系统性缺失时更优
- 适合资源有限但需应对非概率样本的研究者
调查研究面临两大挑战:概率样本成本上升与数据缺失(如无回应或流失),导致推断失效并促使使用便利样本。现有工作尝试用大语言模型(LLMs)通过角色提示生成受访者,常无需标注数据。本文研究更实际的情境:存在部分调查数据时,利用德国纵向选举研究数据,对LLM进行微调以填补自报投票选择,覆盖随机与系统性非响应。比较零样本提示与监督微调,对比树模型(如CatBoost),测试不同便利样本(如学生群体)对泛化能力的影响。结果表明:当数据完全随机缺失时,微调后的LLM性能与表格分类器相当,优于零样本方法;当仅可用有偏便利样本时,微调小规模(3B至8B参数)开源LLM能更准确恢复个体预测与总体分布,常优于零样本及部分表格方法。这表明微调后的LLM是处理非概率样本或系统性缺失数据的可行策略,或可支持仅依赖易获取子群的新调查设计。
原文摘要 · Abstract (English)
Survey researchers face two key challenges: the rising costs of probability samples and missing data (e.g., non-response or attrition), which can undermine inference and increase the use of convenience samples. Recent work explores using large language models (LLMs) to simulate respondents via persona-based prompts, often without labeled data. We study a more practical setting where partial survey responses exist: we fine-tune LLMs on available data to impute self-reported vote choice under both random and systematic nonresponse, using the German Longitudinal Election Study. We compare zero-shot prompting and supervised fine-tuning against tabular classifiers (e.g., CatBoost) and test how different convenience samples (e.g., students) used for fine-tuning affect generalization. Our results show that when data are missing completely at random, fine-tuned LLMs match tabular classifiers but outperform zero-shot approaches. When only biased convenience samples are available, fine-tuning small (3B to 8B) open-source LLMs can recover both individual-level predictions and population-level distributions more accurately than zero-shot and often better than tabular methods. This suggests fine-tuned LLMs offer a promising strategy for researchers working with non-probability samples or systematic missingness, and may enable new survey designs requiring only easily accessible subpopulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。