arXiv:2603.19271cs.CLcs.AI2026-03

教研究者用API方式安全高效地用大模型做内容分析

A Human-Centered Workflow for Using Large Language Models in Content Analysis

  • 将大模型当通用文本处理机,分三步完成标注、摘要和信息提取
  • 每步由人全程设计监督,确保结果可解释且抗幻觉
  • 配套提示库与Python代码,适合社科类研究者快速上手

尽管许多研究者通过聊天界面使用大语言模型(LLMs),但其真正潜力在于通过应用程序接口(API)调用。本文将LLMs视为通用文本处理机器,提出一套完整的工作流程,用于开展三种定性与定量内容分析任务:(1)标注(涵盖质性编码、标签化与文本分类),(2)摘要,(3)信息抽取。该流程强调以人为本,研究人员在每个阶段均参与设计、监督与验证,以保障研究的严谨性与透明度。方法融合了政治学、社会学、计算机科学、心理学及管理学等多学科的实证研究经验,明确了验证程序与最佳实践,以应对LLM的黑箱特性、提示敏感性及幻觉问题。为支持实际应用,论文提供补充材料,包括提示库、Jupyter Notebook格式的Python代码及详细使用说明。

原文摘要 · Abstract (English)

While many researchers use Large Language Models (LLMs) through chat-based access, their real potential lies in leveraging LLMs via application programming interfaces (APIs). This paper conceptualizes LLMs as universal text processing machines and presents a comprehensive workflow for employing LLMs in three qualitative and quantitative content analysis tasks: (1) annotation (an umbrella term for qualitative coding, labeling and text classification), (2) summarization, and (3) information extraction. The workflow is explicitly human-centered. Researchers design, supervise, and validate each stage of the LLM process to ensure rigor and transparency. Our approach synthesizes insights from extensive methodological literature across multiple disciplines: political science, sociology, computer science, psychology, and management. We outline validation procedures and best practices to address key limitations of LLMs, such as their black-box nature, prompt sensitivity, and tendency to hallucinate. To facilitate practical implementation, we provide supplementary materials, including a prompt library and Python code in Jupyter Notebook format, accompanied by detailed usage instructions.

大模型应用内容分析人机协同科研工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。