用大模型代理模拟社科文本分析全流程,提升效率与准确性。
SCALE: Towards Collaborative Content Analysis in Social Science with Large Language Model Agents and Human Intervention
- 构建多智能体系统,模拟人工编码、协作讨论与代码本迭代过程。
- 在真实数据集上达到接近人类专家的分析性能,减少重复劳动。
- 支持专家干预,适合需要高精度的社科研究场景。
内容分析将复杂非结构化文本转化为基于理论的可量化类别。在社会科学中,这一过程通常依赖多轮人工标注、领域专家讨论和基于规则的优化。本文提出SCALE,一种新型多智能体框架,通过大语言模型代理(LLM agents)模拟内容分析的关键阶段:文本编码、协作讨论与动态代码本演化,还原研究人员的反思深度与适应性讨论。通过整合多种人类干预方式,SCALE进一步融合专家知识以提升性能。在多个真实数据集上的广泛评估表明,该方法在各类复杂内容分析任务中达到接近人类水平的表现,为未来社会科学研究提供了创新可能。
原文摘要 · Abstract (English)
Content analysis breaks down complex and unstructured texts into theory-informed numerical categories. Particularly, in social science, this process usually relies on multiple rounds of manual annotation, domain expert discussion, and rule-based refinement. In this paper, we introduce SCALE, a novel multi-agent framework that effectively $\underline{\textbf{S}}$imulates $\underline{\textbf{C}}$ontent $\underline{\textbf{A}}$nalysis via $\underline{\textbf{L}}$arge language model (LLM) ag$\underline{\textbf{E}}$nts. SCALE imitates key phases of content analysis, including text coding, collaborative discussion, and dynamic codebook evolution, capturing the reflective depth and adaptive discussions of human researchers. Furthermore, by integrating diverse modes of human intervention, SCALE is augmented with expert input to further enhance its performance. Extensive evaluations on real-world datasets demonstrate that SCALE achieves human-approximated performance across various complex content analysis tasks, offering an innovative potential for future social science research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。