arXiv:2512.05461cs.CYcs.AI2025-12

教研究者如何评估大模型在社会科学中的不确定性

Knowing Your Uncertainty -- On the application of LLM in social sciences

  • 按任务类型与验证条件构建双重评估框架
  • 明确不同任务下模型输出的可信度边界
  • 适合需要严谨方法的社会科学学者使用

大语言模型正快速融入计算社会科学,但其黑箱训练和推理中的随机性给科学研究带来挑战。本文主张,在社会科学研究中应用大模型需明确评估不确定性,这一要求在社会科学研究与机器学习领域早有共识。我们提出一个统一框架,从任务类型(分类、短文本生成、长文本生成)和验证类型(是否有参考数据或评价标准)两个维度评估模型不确定性。结合计算机科学与社会科学文献,将现有不确定性量化(UQ)方法映射到该分类体系,并为研究者提供实用建议。该框架既保障方法严谨性,也为大模型融入社会科学提供了可操作指南。

原文摘要 · Abstract (English)

Large language models (LLMs) are rapidly being integrated into computational social science research, yet their blackboxed training and designed stochastic elements in inference pose unique challenges for scientific inquiry. This article argues that applying LLMs to social scientific tasks requires explicit assessment of uncertainty-an expectation long established in both quantitative methodology in the social sciences and machine learning. We introduce a unified framework for evaluating LLM uncertainty along two dimensions: the task type (T), which distinguishes between classification, short-form, and long-form generation, and the validation type (V), which captures the availability of reference data or evaluative criteria. Drawing from both computer science and social science literature, we map existing uncertainty quantification (UQ) methods to this T-V typology and offer practical recommendations for researchers. Our framework provides both a methodological safeguard and a practical guide for integrating LLMs into rigorous social science research.

大模型不确定性社会科学方法论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。