用访谈数据训练LLM生成问卷回答,提升量质结合研究的可靠性
Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data
- 用访谈内容指导LLM生成问卷回复,实现定性到定量转化
- 生成回复整体模式相似,但多样性低于真人,尤其在情感表达上
- 精心设计提示词和低温度设置能显著提高生成数据与真实数据的一致性
混合方法研究融合量化与质性数据,但二者结构差异导致整合困难,尤其在测量特征与个体响应模式分析方面。大语言模型(LLMs)通过利用质性数据生成合成问卷回复,提供了新解决方案。本研究以课后项目工作人员访谈和行为锻炼调节量表(BREQ)为案例,考察经访谈引导的LLM能否可靠预测人类问卷回答。结果显示,LLM能捕捉总体响应模式,但变异程度低于人类。引入访谈数据可提升部分模型(如Claude、GPT)的回复多样性;优化提示词设计与采用低温度设置能增强LLM与人类回答的一致性。人口统计信息对一致性影响较小,而访谈内容起主导作用。研究揭示了访谈引导型LLM在弥合质量化方法方面的潜力,也暴露了其在响应多样性、情感理解与心理测量保真度上的局限。未来需改进提示设计、探索偏差缓解策略,并优化模型参数,以提升社会科学研究中生成问卷数据的有效性。
原文摘要 · Abstract (English)
Mixed methods research integrates quantitative and qualitative data but faces challenges in aligning their distinct structures, particularly in examining measurement characteristics and individual response patterns. Advances in large language models (LLMs) offer promising solutions by generating synthetic survey responses informed by qualitative data. This study investigates whether LLMs, guided by personal interviews, can reliably predict human survey responses, using the Behavioral Regulations in Exercise Questionnaire (BREQ) and interviews from after-school program staff as a case study. Results indicate that LLMs capture overall response patterns but exhibit lower variability than humans. Incorporating interview data improves response diversity for some models (e.g., Claude, GPT), while well-crafted prompts and low-temperature settings enhance alignment between LLM and human responses. Demographic information had less impact than interview content on alignment accuracy. These findings underscore the potential of interview-informed LLMs to bridge qualitative and quantitative methodologies while revealing limitations in response variability, emotional interpretation, and psychometric fidelity. Future research should refine prompt design, explore bias mitigation, and optimize model settings to enhance the validity of LLM-generated survey data in social science research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。