arXiv:2505.09724cs.CLcs.AI2025-05被引 1

教研究者用大模型迭代构建文本分类体系,提升分析效率与可靠性。

An AI-Powered Research Assistant in the Lab: A Practical Guide for Text Analysis Through Iterative Collaboration with LLMs

  • 通过人机协作迭代生成文本分类框架
  • 实现对全量数据的高一致性分类标注
  • 适合心理学、社会学等需解析开放文本的研究者

分析开放式回答、新闻标题或社交媒体内容等文本是一项耗时且易受偏见影响的工作。大语言模型(LLMs)可作为高效工具,采用预设(自上而下)或数据驱动(自下而上)的分类体系,在不降低质量的前提下完成分析。本文提供分步教程,展示如何通过研究人员与大模型的持续协作,高效构建、测试并应用针对非结构化数据的分类体系。以参与者提供的个人目标为例,演示如何编写提示语审查数据集、生成生活领域分类体系,通过提示调整和直接修改优化分类,测试分类体系并评估编码者间一致性,最终实现对整批数据的分类,达到高编码者间信度。文章还讨论了大模型在文本分析中的潜力与局限。

原文摘要 · Abstract (English)

Analyzing texts such as open-ended responses, headlines, or social media posts is a time- and labor-intensive process highly susceptible to bias. LLMs are promising tools for text analysis, using either a predefined (top-down) or a data-driven (bottom-up) taxonomy, without sacrificing quality. Here we present a step-by-step tutorial to efficiently develop, test, and apply taxonomies for analyzing unstructured data through an iterative and collaborative process between researchers and LLMs. Using personal goals provided by participants as an example, we demonstrate how to write prompts to review datasets and generate a taxonomy of life domains, evaluate and refine the taxonomy through prompt and direct modifications, test the taxonomy and assess intercoder agreements, and apply the taxonomy to categorize an entire dataset with high intercoder reliability. We discuss the possibilities and limitations of using LLMs for text analysis.

文本分析大模型人机协作分类体系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。