用ChatGPT做法律案例分类,通过新策略实现可靠结果
Assessing the Reliability of Large Language Models for Deductive Qualitative Coding: A Comparative Study of ChatGPT Interventions
- 设计分步任务分解法提升LLM分类一致性
- 最佳方法准确率达77.5%,达实质性一致标准
- 适合需要高效且可靠的定性编码研究者
本研究探讨大型语言模型(特别是ChatGPT)在结构化演绎式定性编码中的应用。现有研究多聚焦归纳编码,本文则关注LLM执行与人工编码体系一致的演绎分类任务。基于比较议程项目(CAP)主编码手册,将美国最高法院案例摘要划分为21个主要政策领域。测试了零样本、少样本、基于定义及一种新型分步任务分解策略,在重复样本上评估性能。采用准确率、F1分数、Cohen's kappa和Krippendorff's alpha等指标,并通过卡方检验和Cramer's V评估构念效度。结果显示干预策略显著影响分类行为,Cramer's V值为0.359至0.613,表明分类模式有中等到强的改变。分步任务分解策略表现最优(准确率=0.775,kappa=0.744,alpha=0.746),达到实质性一致阈值。尽管案例摘要存在语义模糊,ChatGPT在各样本间仍保持稳定,低支持类别的F1分数也较高。结果表明,经定制化干预后,LLM可达到适用于严谨定性编码流程的可靠性水平。
原文摘要 · Abstract (English)
In this study, we investigate the use of large language models (LLMs), specifically ChatGPT, for structured deductive qualitative coding. While most current research emphasizes inductive coding applications, we address the underexplored potential of LLMs to perform deductive classification tasks aligned with established human-coded schemes. Using the Comparative Agendas Project (CAP) Master Codebook, we classified U.S. Supreme Court case summaries into 21 major policy domains. We tested four intervention methods: zero-shot, few-shot, definition-based, and a novel Step-by-Step Task Decomposition strategy, across repeated samples. Performance was evaluated using standard classification metrics (accuracy, F1-score, Cohen's kappa, Krippendorff's alpha), and construct validity was assessed using chi-squared tests and Cramer's V. Chi-squared and effect size analyses confirmed that intervention strategies significantly influenced classification behavior, with Cramer's V values ranging from 0.359 to 0.613, indicating moderate to strong shifts in classification patterns. The Step-by-Step Task Decomposition strategy achieved the strongest reliability (accuracy = 0.775, kappa = 0.744, alpha = 0.746), achieving thresholds for substantial agreement. Despite the semantic ambiguity within case summaries, ChatGPT displayed stable agreement across samples, including high F1 scores in low-support subclasses. These findings demonstrate that with targeted, custom-tailored interventions, LLMs can achieve reliability levels suitable for integration into rigorous qualitative coding workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。