arXiv:2509.24294cs.CLcs.HC2025-09被引 5

用大模型自动构建质性研究理论,省去人工编码瓶颈

LOGOS: LLM-driven End-to-End Grounded Theory Development and Schema Induction for Qualitative Research

  • 用大模型+图推理实现从文本到理论的全流程自动化
  • 在5个数据集上平均达80.4%与专家理论对齐度
  • 适合希望高效开展质性研究的学者或跨领域团队

扎根理论能深入挖掘质性数据洞察,但依赖专家手动编码,难以规模化。现有计算工具或无法完全自动化,或缺乏灵活的编码体系构建能力。我们提出LOGOS,一个端到端框架,可将原始文本全自动转化为结构化、分层化的理论体系。LOGOS融合大模型编码、语义聚类、图推理及一种新颖的迭代优化流程,生成高度可复用的编码手册。为确保公平比较,我们还引入五维评估指标和标准化训练-测试划分协议。在五个不同语料库上,LOGOS持续优于强基线,在复杂数据集上平均达到80.4%与专家构建理论的对齐度。LOGOS展现了在不牺牲理论深度的前提下,推动质性研究民主化与规模化的潜力。

原文摘要 · Abstract (English)

Grounded theory offers deep insights from qualitative data, but its reliance on expert-intensive manual coding presents a major scalability bottleneck. Existing computational tools either fail on full automation or lack flexible schema construction. We introduce LOGOS, a novel, end-to-end framework that fully automates the grounded theory workflow, transforming raw text into a structured, hierarchical theory. LOGOS integrates LLM-driven coding, semantic clustering, graph reasoning, and a novel iterative refinement process to build highly reusable codebooks. To ensure fair comparison, we also introduce a principled 5-dimensional metric and a train-test split protocol for standardized, unbiased evaluation. Across five diverse corpora, LOGOS consistently outperforms strong baselines and achieves a remarkable average $80.4\%$ alignment with an expert-developed schema on complex datasets. LOGOS demonstrates a potential to democratize and scale qualitative research without sacrificing theoretical nuance.

质性研究大模型应用理论构建自动化编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。