用大模型当裁判,自动评估法律文档推荐系统效果
LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
- 让大模型充当评审员,替代人工打分
- 发现传统评分一致性指标易误判,新指标更可靠
- 适合需要高精度评估的法律AI系统研发者
生成式AI兴起使推荐系统评估面临瓶颈,传统指标难以捕捉法律等专业领域中的质量细节。本文探究大模型作为评审员(LLM-as-a-Judge)在检索增强生成系统中的应用可行性,重点解决两个核心问题:何种评价指标能准确反映大模型与人类判断的一致性?如何实现统计上严谨的系统对比?实验表明,传统一致性指标如Krippendorff's alpha在偏态分布下可能误导结果;而Gwet's AC2和秩相关系数更具鲁棒性,适用于评审员筛选;结合贝尼特-霍克伯格校正的威尔科森符号秩检验可提供可靠的系统比较。研究为法律场景下高精度、低成本的自动化评估提供了可扩展框架。
原文摘要 · Abstract (English)
The evaluation bottleneck in recommendation systems has become particularly acute with the rise of Generative AI, where traditional metrics fall short of capturing nuanced quality dimensions that matter in specialized domains like legal research. Can we trust Large Language Models to serve as reliable judges of their own kind? This paper investigates LLM-as-a-Judge as a principled approach to evaluating Retrieval-Augmented Generation systems in legal contexts, where the stakes of recommendation quality are exceptionally high. We tackle two fundamental questions that determine practical viability: which inter-rater reliability metrics best capture the alignment between LLM and human assessments, and how do we conduct statistically sound comparisons between competing systems? Through systematic experimentation, we discover that traditional agreement metrics like Krippendorff's alpha can be misleading in the skewed distributions typical of AI system evaluations. Instead, Gwet's AC2 and rank correlation coefficients emerge as more robust indicators for judge selection, while the Wilcoxon Signed-Rank Test with Benjamini-Hochberg corrections provides the statistical rigor needed for reliable system comparisons. Our findings suggest a path toward scalable, cost-effective evaluation that maintains the precision demanded by legal applications, transforming what was once a human-intensive bottleneck into an automated, yet statistically principled, evaluation framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。