arXiv:2505.24826cs.CLcs.CV2025-05被引 3

新基准LegalEval-Q评估法律文本质量,发现140亿参数后性能趋于稳定。

LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text

  • 基于清晰度、连贯性、术语使用构建文本质量回归模型
  • 49个模型测试显示140亿参数后性能提升仅2.7%
  • 推理型模型优于基础架构,推荐通义千问Qwen3系列

随着大语言模型在法律领域的应用日益广泛,现有评估基准多关注事实准确性,却忽视了清晰度、连贯性和术语使用等语言质量维度。为此,我们提出三步方案:首先,构建基于清晰度、连贯性和术语使用的文本质量回归评估模型;其次,设计专门的法律问题集;最后,采用该框架评估49个大语言模型。分析发现:模型性能在140亿参数时趋于饱和,720亿参数仅带来2.7%的微弱提升;量化与上下文长度等工程选择影响不显著(统计显著性阈值>0.016);推理类模型始终优于基础架构。研究还发布了排名列表与帕累托分析,指出Qwen3系列在成本-性能权衡上表现最优。本工作不仅建立了法律大模型标准化评估协议,也揭示了当前训练数据精炼方法的根本局限。代码与模型已开源:https://github.com/lyxx3rd/LegalEval-Q。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly used in legal applications, current evaluation benchmarks tend to focus mainly on factual accuracy while largely neglecting important linguistic quality aspects such as clarity, coherence, and terminology. To address this gap, we propose three steps: First, we develop a regression model to evaluate the quality of legal texts based on clarity, coherence, and terminology. Second, we create a specialized set of legal questions. Third, we analyze 49 LLMs using this evaluation framework. Our analysis identifies three key findings: First, model quality levels off at 14 billion parameters, with only a marginal improvement of $2.7\%$ noted at 72 billion parameters. Second, engineering choices such as quantization and context length have a negligible impact, as indicated by statistical significance thresholds above 0.016. Third, reasoning models consistently outperform base architectures. A significant outcome of our research is the release of a ranking list and Pareto analysis, which highlight the Qwen3 series as the optimal choice for cost-performance tradeoffs. This work not only establishes standardized evaluation protocols for legal LLMs but also uncovers fundamental limitations in current training data refinement approaches. Code and models are available at: https://github.com/lyxx3rd/LegalEval-Q.

法律AI大模型评估质量评测通义千问

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。