arXiv:2605.13412cs.CLcs.AI2026-05中稿 · the 20th Linguisti…

用大模型标注丹麦难民判决中的可信度评估,发现效果不错但误差不一致。

LLMs as annotators of credibility assessment in Danish asylum decisions: evaluating classification performance and errors beyond aggregated metrics

论文配图:LLMs as annotators of credibility assessment in Danish asylum decisions: evaluating classification performance and errors beyond aggregated metrics
图 1 · 摘自论文原文
  • 用21个开源模型+30种提示词测试大模型在难民判决文本中识别可信度的能力
  • 顶级模型在零样本和少样本下准确率超85%,但错误模式不一致且与人类信心相关
  • 提出新数据集RAB-Cred,适合法律文本、小语种和可信度分析研究者使用

现成的大语言模型(LLMs)正被用于自动化文本标注,但在低资源语言和需要专家判断的领域中其有效性仍不清楚。本文研究了大模型在一项新型法律自然语言处理任务中的应用:识别丹麦难民判决书中可信度评估的存在性与情感倾向。我们构建了RAB-Cred数据集,包含高质量专家标注的丹麦文本,附带标注者置信度与案件结果等元数据。对21个开源模型和30种系统-用户提示组合进行了基准测试,系统评估了模型与提示选择对零样本和少样本分类的影响。深入分析了表现最佳模型与提示的错误,考察了错误在不同大模型间的稳定性、类别间混淆情况、与人工置信度的相关性,以及样本层面的难度与错误严重性。结果表明大模型具备低成本标注难民判决的潜力,但其标注行为存在不一致性和局限性,需避免依赖单一模型的预测。RAB-Cred数据集与代码已公开于https://github.com/glhr/RAB-Cred。

原文摘要 · Abstract (English)

Off-the-shelf large language models (LLMs) are increasingly used to automate text annotation, yet their effectiveness remains underexplored for underrepresented languages and specialized domains where the class definition requires subtle expert understanding. We investigate LLM-based annotation for a novel legal NLP task: identifying the presence and sentiment of credibility assessments in asylum decision texts. We introduce RAB-Cred, a Danish text classification dataset featuring high-quality, expert annotations and valuable metadata such as annotator confidence and asylum case outcome. We benchmark 21 open-weight models and 30 system-user prompt combinations for this task, and systematically evaluate the effect of model and prompt choice for zero-shot and few-shot classification. We zoom in on the errors made by top-performing models and prompts, investigating error consistency across LLMs, inter-class confusion, correlation with human confidence and sample-wise difficulty and severity of LLM mistakes. Our results confirm the potential of LLMs for cost-effective labeling of asylum decisions, but highlight the imperfect and inconsistent nature of LLM annotators, and the need to look beyond the predictions of a single, arbitrarily chosen model. The RAB-Cred dataset and code are available at https://github.com/glhr/RAB-Cred

法律NLP大模型标注丹麦语难民决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。