arXiv:2504.03822cs.CYcs.AI2025-04被引 4

教社科研究者用大模型做科学推断,避免盲目信任输出

Arti-"fickle" Intelligence: Using LLMs as a Tool for Inference in the Political and Social Sciences

  • 强调用大模型推断人类行为时需聚焦科学目标
  • 提出验证模型输出成败的标准框架
  • 适合关注实证研究可信度的社科学者

生成式大语言模型(LLMs)在政治与社会科学研究中具有巨大潜力。但其真正价值在于推动对真实人类行为与关切的理解。为促进其科学应用,本文强调研究者应始终以科学推断为目标。通过模型输出验证这一具体案例,我们探讨了使用大模型进行科学推断所面临的挑战与机遇,并提出一套用于判断大模型在特定任务中成功或失败的标准。这些标准有助于从模型表现中做出可靠推断。最后,本文讨论了这一方法论转变如何促进社会科学领域对大模型及其应用的共享知识积累。

原文摘要 · Abstract (English)

Generative large language models (LLMs) are incredibly useful, versatile, and promising tools. However, they will be of most use to political and social science researchers when they are used in a way that advances understanding about real human behaviors and concerns. To promote the scientific use of LLMs, we suggest that researchers in the political and social sciences need to remain focused on the scientific goal of inference. To this end, we discuss the challenges and opportunities related to scientific inference with LLMs, using validation of model output as an illustrative case for discussion. We propose a set of guidelines related to establishing the failure and success of LLMs when completing particular tasks, and discuss how we can make inferences from these observations. We conclude with a discussion of how this refocus will improve the accumulation of shared scientific knowledge about these tools and their uses in the social sciences.

大模型社会科学推断可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。