arXiv:2509.05830cs.LGcs.CY2025-09EMNLP被引 36

用真实实验数据微调大模型,提升社会实验预测准确率

Finetuning LLMs for Human Behavior Prediction in Social Science Experiments

  • 直接在个体响应数据上微调,提升跨领域预测能力
  • 新研究中预测匹配度比基线高26%,优于GPT-4o13%
  • 适合用于实验假设筛选的社会科学模拟场景

大型语言模型(LLM)为模拟社会科学实验结果提供了强大机会。本文证明,将LLM直接在过往实验的个体级响应数据上进行微调,可显著提升多种社会科学领域的模拟准确性。我们通过自动化流程构建了SocSci210数据集,包含来自210个开源社会科学研究、400,491名参与者共计290万条响应。通过微调,实现多层级泛化:在完全未见的研究中,最强模型Socrates-Qwen-14B对多样化结果问题的预测分布与人类响应的匹配度比其基础模型Qwen2.5-14B高出26%,较GPT-4o提升13%;在某一研究的子条件上微调后,对新未见条件的泛化能力提升71%。由于SocSci210包含丰富人口统计信息,通过微调使偏差指标(人口平等差异)降低10.6%。鉴于社会科学研究常产生丰富且主题特定的数据,本研究表明在这些数据上微调可实现更精准的实验模拟,助力假设筛选。数据、模型与微调代码已开源,地址:stanfordhci.github.io/socrates。

原文摘要 · Abstract (English)

Large language models (LLMs) offer a powerful opportunity to simulate the results of social science experiments. In this work, we demonstrate that finetuning LLMs directly on individual-level responses from past experiments meaningfully improves the accuracy of such simulations across diverse social science domains. We construct SocSci210 via an automatic pipeline, a dataset comprising 2.9 million responses from 400,491 participants in 210 open-source social science experiments. Through finetuning, we achieve multiple levels of generalization. In completely unseen studies, our strongest model, Socrates-Qwen-14B, produces predictions that are 26% more aligned with distributions of human responses to diverse outcome questions under varying conditions relative to its base model (Qwen2.5-14B), outperforming GPT-4o by 13%. By finetuning on a subset of conditions in a study, generalization to new unseen conditions is particularly robust, improving by 71%. Since SocSci210 contains rich demographic information, we reduce demographic parity difference, a measure of bias, by 10.6% through finetuning. Because social sciences routinely generate rich, topic-specific datasets, our findings indicate that finetuning on such data could enable more accurate simulations for experimental hypothesis screening. We release our data, models and finetuning code at stanfordhci.github.io/socrates.

大模型微调社会实验行为预测数据模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。