用迁移学习复用调查数据,生成高精度模拟回答
Survey Transfer Learning: Recycling Data with Silicon Responses
- 基于政治行为理论,利用共享人口变量跨调查迁移知识
- 在CES和ANES数据上准确率高达93%,敏感问题更优
- 比大模型更省钱透明,适合社会科学研究与民调
随着研究人员越来越多地使用大语言模型(LLMs)生成合成调查数据,对替代性AI范式的关注却较少,尤其考虑到LLMs的环境成本。本文提出调查迁移学习(Survey Transfer Learning, STL),将计算机科学中的迁移学习范式引入调查研究,以复用已有调查数据并生成具有实证基础的硅基响应。受美国政治行为理论启发,STL利用在极化美国语境下具有高预测力的共享人口变量,在合作选举研究(CES)2020上预训练神经网络,冻结早期层以保留学习结构,并在2020年美国全国选举研究(ANES)上微调顶层。该方法在生成CES 2022及预留的ANES 2020数据的硅基响应时,准确率最高达93%。结果表明,STL在敏感测量如种族怨恨方面优于LLMs。相较于成本高昂且不透明的LLM模拟样本,STL能生成高个体层级准确度的实证响应,有望缓解社会科学与民调行业的关键挑战。
原文摘要 · Abstract (English)
As researchers increasingly turn to large language models (LLMs) to generate synthetic survey data, less attention has been paid to alternative AI paradigms given environmental costs of LLMs. This paper introduces Survey Transfer Learning (STL), which develops transfer learning paradigms from computer science for survey research to recycle existing survey data and generate empirically grounded silicon responses. Inspired by political behavior theory, STL leverages shared demographic variables with high predictive power in a polarized American context to transfer knowledge across surveys. Using a neural network pre-trained on the Cooperative Election Study (CES) 2020, freezing early layers to preserve learned structure, and fine-tuning top layers on the American National Election Studies (ANES) 2020, STL generates silicon responses CES 2022 and in held-out ANES 2020 data with accuracy rates of up to 93 percent. Results show that STL outperforms LLMs, especially on sensitive measures such as racial resentment. While LLMs silicon samples are costly and opaque, STL generates empirically grounded silicon responses with high individual-level accuracy, potentially helping to mitigate key challenges in social science and the polling industry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。