用真实推特数据评估大模型回复的语义偏差,提升社科研究可信度
Next Reply Prediction X Dataset: Linguistic Discrepancies in Naively Generated Content
- 基于真实推特历史构建对话预测任务,检验大模型生成内容差异
- 发现无约束提示的大模型输出存在显著语言风格偏差
- 适合关注大模型用于社会科学研究的学者与伦理审查者
大型语言模型(LLMs)在社会科学实验中作为人类参与者的替代品,虽具可扩展性和低成本优势,但其“盲目”使用——即未施加行为约束的提示生成——会引入显著的语言差异,影响研究有效性。本文通过在真实X(原Twitter)数据上构建历史条件下的回复预测任务,创建了一个用于评估大模型语言输出与人类内容一致性的新数据集。我们采用风格与内容指标分析这些差异,为研究者提供量化框架以评估合成数据的质量与真实性。研究揭示了需采用更复杂的提示技术及专用数据集,以确保大模型生成内容能准确反映人类交流的复杂语言模式,从而提高计算社会科学研究的有效性。
原文摘要 · Abstract (English)
The increasing use of Large Language Models (LLMs) as proxies for human participants in social science research presents a promising, yet methodologically risky, paradigm shift. While LLMs offer scalability and cost-efficiency, their "naive" application, where they are prompted to generate content without explicit behavioral constraints, introduces significant linguistic discrepancies that challenge the validity of research findings. This paper addresses these limitations by introducing a novel, history-conditioned reply prediction task on authentic X (formerly Twitter) data, to create a dataset designed to evaluate the linguistic output of LLMs against human-generated content. We analyze these discrepancies using stylistic and content-based metrics, providing a quantitative framework for researchers to assess the quality and authenticity of synthetic data. Our findings highlight the need for more sophisticated prompting techniques and specialized datasets to ensure that LLM-generated content accurately reflects the complex linguistic patterns of human communication, thereby improving the validity of computational social science studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。