arXiv:2607.07915cs.CYcs.CL2026-07被引 1

检验大模型在社科研究中的可靠性,揭示方法论风险与验证规范缺失

Validating LLMs in social science: Epistemic threats and emerging norms

论文配图:Validating LLMs in social science: Epistemic threats and emerging norms
图 1 · 摘自论文原文
  • 分析八本顶刊论文,发现大模型生成数据常为核心分析依据
  • 多数研究缺乏系统验证,存在偏见与幻觉等效度威胁
  • 提出互补验证策略,推动社科领域大模型使用新规范

大语言模型正重塑社会科学方法论。研究人员越来越多地通过提示词让语言模型生成社会概念的量化测量,如数据标注或模拟调查回应。然而,大模型存在偏见、幻觉及跨情境脆弱性等方法论挑战,其对研究有效性的威胁尚不明确。当前应对这些挑战的标准实践和规范仍处于形成中。本文收集并系统分析了八本旗舰社会科学期刊中使用大模型作为测量工具的论文全集,发现大模型生成的测量结果常在实证分析中起核心作用,但验证实践普遍不一致且有限。研究提出互补性的稳健验证策略,指向更完善的规范与标准,以提升大模型在社会科学中的可信度。

原文摘要 · Abstract (English)

Large language models (LLMs) are reshaping social science methodology. Researchers increasingly prompt language models to generate quantitative measurements of social concepts, for example labeling data or simulating survey responses. Yet LLMs pose methodological challenges including bias, hallucination, and brittleness across contexts, with unclear threats to validity. Standard practices and norms for addressing these challenges are still emerging. We collect and systematically analyze validation practices in a comprehensive corpus of papers from eight flagship social science journals that use LLMs as measurement instruments. We find that LLM-generated measurements frequently play a central role in empirical analyses, yet validation practices are inconsistent and limited. We outline complementary strategies for more robust validation, pointing toward better norms and standards around the use of LLMs in social science.

大模型社会科学研究有效性验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。