arXiv:2605.29025cs.AIcs.CY2026-05

用模型分歧检测公共评论分析中的深层解读复杂性。

When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis

论文配图:When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
图 1 · 摘自论文原文
  • 将多模型分歧视为解读复杂性的诊断信号,引导人工审核
  • 4个LLM对1260条评论的分类差异大于单模型提示变化
  • 适合政策评估、司法审查等需审慎解读的场景

联邦机构正使用大语言模型(LLMs)对公众意见文档进行分类,模型对文本的组织方式直接影响政策制定者所见内容及哪些论点被记录。标准评估以小规模验证集上的立场准确率为锚点,无法识别不同模型对同一公众输入产生实质性分类差异的情况。本文提出一种解释性审计流程,将多模型分歧视为解读复杂性的指示,并引导人工审查聚焦于真正模糊的公众输入。在四个LLM对1260条美国农业部(USDA)文件的意见分析中,发现模型间主题分歧超过单模型提示变异;专家评分标准虽抑制了表面分歧,但未根本解决深层解读矛盾。通过对分层抽样的40条意见进行两阶段标注研究,四个LLM与一人类标注员独立标注后又查看他人结果进行修订。修订行为在不同标注者间差异明显,人类标注员的修订常引入集合输出中不存在的新框架。我们主张,基于分歧的评估应作为大模型辅助解释性编码中准确率指标的必要补充。

原文摘要 · Abstract (English)

Federal agencies are deploying large language models (LLMs) to categorize public comment corpora, where the model's organization of the record shapes what policymakers see and which arguments register. Standard evaluation, anchored on stance accuracy against a small validated set, cannot detect when different models produce materially different categorizations of the same public input. We propose an Interpretive Audit Pipeline that treats multi-model disagreement as diagnostic of interpretive complexity and directs human review toward genuinely ambiguous public input. Analyzing 1,260 public comments on a federal USDA docket across four LLMs, we find that inter-model thematic divergence exceeds within-model prompt variation, and that an expert rubric suppresses deep interpretive disagreement without resolving it. In a two-stage labeling study on a stratified 40-comment subsample, four LLMs and a human annotator labeled independently and then revised after seeing the others' labels. Revision behavior varied across labelers, and the human annotator's revisions frequently introduced framings absent from the ensemble's collective output. We argue disagreement-based evaluation is a necessary complement to accuracy metrics for LLM-assisted interpretive coding.

LLM评估公共政策模型分歧解释性审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。