大模型在民族志文本标注中表现不佳,难替代人工。
Large language models struggle with ethnographic text annotation
- 测试7个顶尖大模型在567段文本上标注121个仪式特征
- 准确率远低于可靠自动化标准,长文本和模糊概念更难处理
- 即使人类共识高,模型仍落后,适合需深度文化理解的研究
大型语言模型(LLMs)在自动化文本标注方面展现出潜力,有望加速跨文化研究,通过从民族志文本中提取结构化数据。我们评估了7个最先进的LLMs在567段民族志摘录中对121个仪式特征的标注能力。结果表明,性能有限,远低于可靠自动化标注所需水平。长文本、需序次判断的特征以及模糊概念尤其难以处理。人类编码者间的一致性设定了LLM准确率的近似上限:人类难以达成一致的特征,对模型同样困难。即便在人类一致性高的特征上,模型表现仍显著低于人类。研究结果表明,目前大模型尚无法在民族志标注中取代人类专家。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise for automated text annotation, raising hopes that they might accelerate cross-cultural research by extracting structured data from ethnographic texts. We evaluated 7 state-of-the-art LLMs on their ability to annotate 121 ritual features across 567 ethnographic excerpts. Performance was limited, falling well below levels required for reliable automated annotation. Longer texts, features requiring ordinal distinctions, and ambiguous constructs proved particularly difficult. Human inter-coder reliability set an approximate ceiling on LLM accuracy: features that human coders found difficult to agree upon were also difficult for LLMs. Yet even on features where humans reliably agreed, models fell short of human performance. Our findings suggest that LLMs cannot yet substitute for human expertise in ethnographic annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。