用置信度与多样性评估大模型在定性编码中的可靠性,减少人工审核量。
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
- 通过置信度和模型多样性构建双信号评估框架
- 可自动接受35%内容,错误率低于5%,人工工作量减65%
- 适合需要大规模定性分析且人类专家意见不一的场景
大语言模型可实现大规模定性编码,但可靠性评估仍具挑战,因人类专家常无法达成一致。本文研究置信度-多样性校准作为可访问编码任务的质量评估框架,针对当前表现良好但存在过度自信问题的LLM。分析来自八个顶尖LLM在十个类别上的5,680个编码决策发现,平均自信心与模型间一致性高度相关(皮尔逊相关系数r=0.82)。引入归一化香农熵衡量的模型多样性后,双信号可几乎完全解释一致性(决定系数R²=0.979),尽管该高预测力可能源于当前任务对LLM而言过于简单。该框架支持三阶段自动化流程,可自动接受35%片段,错误率低于5%,人工工作量减少65%。跨领域验证表明其具备可迁移性(克朗巴赫系数kappa提升0.20至0.78)。该方法为AI判断校准建立方法基础,未来潜力或在于更复杂场景中,大模型相较人类认知局限展现比较优势。
原文摘要 · Abstract (English)
LLMs enable qualitative coding at large scale, but assessing reliability remains challenging where human experts seldom agree. We investigate confidence-diversity calibration as a quality assessment framework for accessible coding tasks where LLMs already demonstrate strong performance but exhibit overconfidence. Analysing 5,680 coding decisions from eight state-of-the-art LLMs across ten categories, we find that mean self-confidence tracks inter-model agreement closely (Pearson r=0.82). Adding model diversity quantified as normalised Shannon entropy produces a dual signal explaining agreement almost completely (R-squared=0.979), though this high predictive power likely reflects task simplicity for current LLMs. The framework enables a three-tier workflow auto-accepting 35 percent of segments with less than 5 percent error, cutting manual effort by 65 percent. Cross-domain validation confirms transferability (kappa improvements of 0.20 to 0.78). While establishing a methodological foundation for AI judgement calibration, the true potential likely lies in more challenging scenarios where LLMs may demonstrate comparative advantages over human cognitive limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。