arXiv:2607.28890cs.HCcs.AI2026-07

人类共识不等于真实质量,机器与人编码谁优谁劣需盲评验证

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

  • 用盲评专家对比人类与LLM的编码结果,打破以一致率为金标准的惯性
  • 人类间一致率0.52,人机间仅0.30,但盲评中两者被认可率几乎持平(51.5%对48.5%)
  • 发现人类集体共识存在偏见,部分代码反而更倾向机器解读,适合评估自动化系统

当前多数LLM辅助定性编码评估以与人工编码的一致性为指标,隐含假设人工编码即真实标准。本研究通过实证表明该假设在一致性指标无法察觉的层面失效。五种LLM系统与三位训练过的编码员独立使用72项层级编码表处理来自K-12 AI平台的2,560条教育者消息。除常规一致性分析外,一位独立领域专家对855组编码对进行盲评,不区分来源,对人类与机器编码一视同仁。两种评估方式在两个方向均出现分歧:人类-机器一致率(均值Jaccard 0.30)显著低于人类-人类一致率(0.52),按传统标准应判定机器表现差;但盲评显示二者被接受率相近(51.5% vs. 48.5%,p = 0.537),Bradley-Terry排名还将两个LLM置于三个编码员之上。对于若干实质性编码,人类共识体现共享偏见,而盲评专家更倾向于机器解释。因此,仅靠一致性评价不足以支撑自动化决策,本研究提出可迁移的验证流程与代码级分工框架。

原文摘要 · Abstract (English)

Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate. This study provides empirical evidence that the presumption fails in ways agreement metrics cannot detect. Five LLM systems and three trained human coders independently applied a 72-item hierarchical codebook to 2,560 educator messages from a K-12 AI platform. Beyond conventional agreement analysis, an independent domain expert judged 855 pairwise comparisons of code sets blind to source, treating human and machine sources symmetrically. The two evaluation approaches diverge in both directions. Human-LLM agreement (mean Jaccard 0.30) falls well below human-human agreement (0.52), which standard practice would read as inferior LLM coding, yet the blind verifier preferred human and LLM coding at indistinguishable rates (51.5% vs. 48.5%, p = 0.537), and a Bradley-Terry ranking placed two LLMs above two of three human coders. For several substantive codes, human consensus encoded shared bias that the verifier rejected in favor of the LLM interpretation. Agreement-based evaluation is therefore insufficient for automation decisions, and the study demonstrates a transferable verification protocol and a code-level division-of-labor framework.

定性编码人类偏见盲评验证大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。