微调小样本人类数据可让大模型更像真人,但仍不能替代真实实验。
Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?
- 用少量真人问卷数据微调大模型,提升多样性与行为一致性。
- 微调后模型在信念与行为匹配上显著优于原始模型。
- 虽有改善,但无法复现真实研究的回归系数,不适合做统计推断。
关于大语言模型(LLMs)能否替代人类参与调查与实验研究,学界仍有争议。尽管营销与心理学领域已尝试用LLM进行模拟,但越来越多证据表明,其行为与真实人类存在系统性偏差:多样性不足、少数群体对齐差、组内方差小,且言与行不一致。本研究聚焦一个关键问题:是否可用小规模真人数据(如试点研究所得)微调LLM,以缓解上述问题并生成真实模拟结果?基于信息披露行为实验,我们从分布差异、子群对齐、信念-行动一致性及回归系数恢复等维度对比了真人与LLM生成数据。结果显示,小样本微调显著提升了异质性、对齐度与信念-行动一致性;但即使表现最佳的微调模型,也无法复现原研究的回归系数,表明其仍不适合作为正式推断分析的人类替代数据。
原文摘要 · Abstract (English)
There is ongoing debate about whether large language models (LLMs) can serve as substitutes for human participants in survey and experimental research. While recent work in fields such as marketing and psychology has explored the potential of LLM-based simulation, a growing body of evidence cautions against this practice: LLMs often fail to align with real human behavior, exhibiting limited diversity, systematic misalignment for minority subgroups, insufficient within-group variance, and discrepancies between stated beliefs and actions. This study examines an important and distinct question in this domain: whether fine-tuning on a small subset of human survey data, such as that obtainable from a pilot study, can mitigate these issues and yield realistic simulated outcomes. Using a behavioral experiment on information disclosure, we compare human and LLM-generated responses across multiple dimensions, including distributional divergence, subgroup alignment, belief-action coherence, and the recovery of regression coefficients. We find that fine-tuning on small human samples substantially improves heterogeneity, alignment, and belief-action coherence relative to the base model. However, even the best-performing fine-tuned models fail to reproduce the regression coefficients of the original study, suggesting that LLM-generated data remain unsuitable for replacing human participants in formal inferential analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。