arXiv:2606.26541cs.LGcs.CY2026-06被引 1

46个大模型在人道数据编码上表现堪比专家,但需谨慎使用。

Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

  • 用结构化提示和推理配置提升编码可靠性
  • 多模型达到与人类专家相当的信度水平
  • 适合需要规模化分析且重视数据治理的救援机构

人道主义组织依赖受影响人群的数据来制定响应策略,但其价值取决于对复杂需求描述的及时、一致解读。当前机构常缺乏人力、时间与专业能力进行大规模分析。本文通过150份高保真合成人道主义语料,对比46个大语言模型与人类黄金标准的表现。评估结合了评分者间一致性测试、Krippendorff's alpha系数、差异性分析(区分正确、近似正确与错误编码),以及针对歧视、复杂需求层级、非标准表达等特定标准的定性评估。结果显示,多个大模型在结构化提示与推理增强配置下,可实现与经验人类编码员相当的推断编码可靠性。然而,整体信度指标不足以决定部署。各模型在识别间接表达的需求、未预设类别中的需求及安全与歧视等保护性议题方面表现差异显著。研究建议:大模型可显著扩展人道分析能力,但不能替代人类判断;应采用结构化编码手册、推理增强模型、关注主题特异性表现,并建立分层监督机制,尤其聚焦误判后果严重的领域。对于敏感数据,自托管基础设施上的开源权重模型或为兼顾分析规模与数据治理的可行路径。

原文摘要 · Abstract (English)

Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanced accounts of need. Humanitarian organizations often lack the staff, time, and specialist expertise required to analyze this information at scale. Large language models (LLMs) may expand this capacity, but their reliability for coding qualitative humanitarian data has not been directly established. This benchmark study compares 46 LLMs to a human Gold Standard using 150 high-fidelity synthetic humanitarian transcripts. Evaluation combined inter-rater reliability testing with Krippendorff's alpha, discrepancy analysis distinguishing correct, near-correct, and incorrect codes, and qualitative assessment across humanitarian-specific criteria including discrimination, complex needs hierarchies, and non-standard communication styles. The authors find that multiple LLMs can perform deductive coding at reliability levels comparable to experienced human coders, especially when structured prompts and reasoning-enabled configurations are used. At the same time, aggregate reliability metrics alone are insufficient for deployment decisions. Models varied in recognizing needs expressed indirectly, needs outside predefined categories, and protection-relevant concerns such as physical safety and discrimination. These findings suggest that LLMs can materially expand humanitarian analytical capacity, but not as substitutes for human judgment. Appropriate use requires structured codebooks, reasoning-enabled models, attention to theme-specific performance, and tiered oversight focused on categories where miscoding would have the greatest programmatic consequences. For sensitive humanitarian data, open-weights models deployed on self-hosted infrastructure may offer a viable path for combining analytical scalability with stronger data governance.

大模型人道主义数据编码可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。