用大模型自动提取文本主题,让大规模语料分析更精准高效
HICode: Hierarchical Inductive Coding with LLMs
- 基于定性研究思路,分两步自动生成标签并构建主题层级
- 在三个数据集上与人工主题对齐度高,经人机评估验证稳定
- 适用于法律、社会等需要深度语义分析的场景
尽管细粒度语料分析应用广泛,研究者仍依赖难以扩展的人工标注,或难以控制的统计工具如主题建模。我们提出,大语言模型有潜力将原本需手动完成的细致分析拓展至大规模文本。为此,受定性研究启发,我们设计了HICode——一种两阶段流程:首先从分析数据中归纳生成标签,再对标签进行层级聚类以揭示潜在主题。我们在三个不同数据集上验证该方法,通过衡量与人工构建主题的一致性,并结合自动化与人工评估证明其稳健性。最后,以美国持续的阿片类药物危机诉讼文件为案例,揭示制药公司采取的激进营销策略,展示了HICode在大规模数据中实现细致分析的潜力。
原文摘要 · Abstract (English)
Despite numerous applications for fine-grained corpus analysis, researchers continue to rely on manual labeling, which does not scale, or statistical tools like topic modeling, which are difficult to control. We propose that LLMs have the potential to scale the nuanced analyses that researchers typically conduct manually to large text corpora. To this effect, inspired by qualitative research methods, we develop HICode, a two-part pipeline that first inductively generates labels directly from analysis data and then hierarchically clusters them to surface emergent themes. We validate this approach across three diverse datasets by measuring alignment with human-constructed themes and demonstrating its robustness through automated and human evaluations. Finally, we conduct a case study of litigation documents related to the ongoing opioid crisis in the U.S., revealing aggressive marketing strategies employed by pharmaceutical companies and demonstrating HICode's potential for facilitating nuanced analyses in large-scale data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。