为社科研究中的大模型应用提供深度与自主性评估框架
Depth and Autonomy: A Framework for Evaluating LLM Applications in Social Science Research
- 从解释深度和自主程度双维度分类大模型应用
- 建议低自主性+受控深度提升研究可靠性
- 适合关注可解释性的社科定量与质性研究者
大型语言模型(LLMs)在社会科学领域日益普及,但其应用仍面临解释偏差、可靠性低和审计困难等挑战。本文提出一个二维评估框架,将大模型应用置于解释深度与自主程度两个维度上,帮助研究人员分类并设计更可靠的使用方式。基于Web of Science收录的全部使用大模型作为工具而非研究对象的社会科学论文,我们梳理了当前研究现状。该框架主张不赋予模型过度自由,而是像指导本科生研究助理一样,将任务分解为可管理的模块。通过保持较低的自主性,并仅在必要且受监督的情况下适度提升解释深度,可在保障透明性与可靠性的前提下,有效利用大模型优势。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly utilized by researchers across a wide range of domains, and qualitative social science is no exception; however, this adoption faces persistent challenges, including interpretive bias, low reliability, and weak auditability. We introduce a framework that situates LLM usage along two dimensions, interpretive depth and autonomy, thereby offering a straightforward way to classify LLM applications in qualitative research and to derive practical design recommendations. We present the state of the literature with respect to these two dimensions, based on all published social science papers available on Web of Science that use LLMs as a tool and not strictly as the subject of study. Rather than granting models expansive freedom, our approach encourages researchers to decompose tasks into manageable segments, much as they would when delegating work to capable undergraduate research assistants. By maintaining low levels of autonomy and selectively increasing interpretive depth only where warranted and under supervision, one can plausibly reap the benefits of LLMs while preserving transparency and reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。