arXiv:2608.27746cs.IR2026-08

构建首个巴西葡萄牙语法律文档检索数据集,评估大模型判别相关性效果。

NormasTCU --- A Brazilian Portuguese IR Dataset and an Evaluation of LLM-as-a-Judge for Relevance Assessment

论文配图:NormasTCU --- A Brazilian Portuguese IR Dataset and an Evaluation of LLM-as-a-Judge for Relevance Assessment
图 1 · 摘自论文原文
  • 构建14,469篇法律文档的巴西葡语检索数据集,含46个查询与3,048条人工标注。
  • 大模型评分存在系统性偏高,与人工标注一致性为中等(Cohen's kappa 0.32-0.53)。
  • 在nDCG@10和MRR指标下,大模型排名与人工排名高度一致,适合规模化评估。

葡萄牙语信息检索缺乏公开数据集,专业文献的相关性评估成本高昂。尽管大语言模型(LLMs)越来越多地用于相关性判断,其在非英语专业领域的可靠性仍不明确。本文引入NormasTCU(https://huggingface.co/datasets/LeandroRibeiro/NormasTCU),一个包含14,469篇巴西葡萄牙语法律文档、46个查询及3,048条人工标注(覆盖812对查询-文档)的数据集。利用该数据集,我们通过两种提示策略评估了三个模型作为评判者的相关性判断能力,并将基于大模型生成的qrels与15个IR系统的排序结果进行比较。结果显示,大模型普遍存在正向评分偏差(0-2分制下平均绝对误差0.46–0.66)。在成对一致性上,与人工标注仅达中等水平,Cohen's kappa为0.32至0.53。尽管如此,大模型生成的排名在nDCG@10和MRR指标上与参考排名高度一致(肯德尔等级相关系数≥0.90,但置信区间未全部超过此阈值),而在P@10和R@10上表现较弱。值得注意的是,大模型排名有时甚至比个别专家标注更接近真实排序。实践上,建议在使用nDCG或MRR等秩相关指标时,借助大模型实现巴西葡语专业文献的可扩展相关性评估,但在依赖精确率或召回率时应谨慎使用。

原文摘要 · Abstract (English)

Portuguese Information Retrieval (IR) lacks public datasets, and relevance assessment for specialized collections remains costly. While Large Language Models (LLMs) increasingly support relevance assessment, their reliability in non-English specialized domains remains unclear. We introduce NormasTCU (https://huggingface.co/datasets/LeandroRibeiro/NormasTCU), a Brazilian Portuguese IR dataset with 14,469 legal documents, 46 queries, and 3,048 human judgments over 812 query-document pairs. Using NormasTCU, we evaluated LLM-as-a-judge for relevance assessment by prompting three models with two prompt techniques to grade these pairs. We then compared the rankings of 15 IR systems derived from LLM-generated and human reference qrels. LLMs consistently showed a positive scoring bias (mean absolute error: 0.46--0.66 on a 0-2 scale). Furthermore, pair-level agreement with human judgments achieved only fair to moderate levels, with Cohen's kappa ranging from 0.32 to 0.53. Despite this bias, LLM-generated judgments often yielded highly similar system rankings for nDCG@10 and MRR (observed Kendall's tau greater than or equal 0.90, although the bootstrap confidence intervals did not always remain above this threshold), but were less reliable for P@10 and R@10. Notably, LLM-based rankings were sometimes more strongly correlated with the reference ranking than individual human annotations were. As a practical implication, our results suggest that LLMs could effectively support scalable relevance assessment in specialized Portuguese corpora when evaluated using nDCG or MRR (rank-aware metrics), but they should be avoided when relying on precision or recall.

信息检索大模型评估多语言法律文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。