教社会科学家用大模型做文本标注,同时提醒潜在风险
Navigating the Risks of Using Large Language Models for Text Annotation in Social Science Research
- 提出一套结合大模型的文本标注框架,支持主判或辅助
- 评估了大模型在分类任务中的有效性与可靠性问题
- 适合关注方法论严谨性的社科研究者参考
大语言模型(LLMs)有望革新计算社会科学,特别是在自动化文本分析方面。本文系统评估了在社会运动研究中使用LLMs进行文本分类的潜力与风险。我们提出一个框架,帮助社会科学家将LLMs用于文本标注,既可作为主要编码决策者,也可作为辅助工具。该框架提供工具以优化提示词设计,并系统检验和报告LLMs在效度、信度、可重复性和透明性方面的表现。此外,我们探讨了其带来的认识论风险,并最终给出若干实用指南,以及更有效地传达这些风险的研究建议。
原文摘要 · Abstract (English)
Large language models (LLMs) have the potential to revolutionize computational social science, particularly in automated textual analysis. In this paper, we conduct a systematic evaluation of the promises and risks associated with using LLMs for text classification tasks, using social movement studies as an example. We propose a framework for social scientists to incorporate LLMs into text annotation, either as the primary coding decision-maker or as a coding assistant. This framework offers researchers tools to develop the potential best-performing prompt, and to systematically examine and report the validity and reliability of LLMs as a methodological tool. Additionally, we evaluate and discuss its epistemic risks associated with validity, reliability, replicability, and transparency. We conclude with several practical guidelines for using LLMs in text annotation tasks and offer recommendations for more effectively communicating epistemic risks in research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。