arXiv:2509.16093cs.CLcs.AI2025-09EMNLP被引 11

拆解答案质量:精准与覆盖分开评估,更懂专家判断。

Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses

  • 用自动提取的评分标准,分精度和覆盖率两维度评估答案
  • 与专家打分相关性达0.78,远超传统方法的0.12~0.48
  • 可揭示模型优劣权衡,适合法律医疗等高风险领域

在法律、医疗等高风险领域,长文本回答的评估仍面临根本挑战。标准指标如BLEU和ROUGE无法捕捉语义正确性,现有基于大模型的评估器常将复杂质量维度压缩为单一分数。本文提出DeCE——一种分解式大模型评估框架,将答案质量分为精确性(事实准确性和相关性)与召回率(必要概念覆盖度),并从标准答案要求中自动提取实例特定的评判标准。DeCE具备模型无关性和领域通用性,无需预设分类体系或人工制定评分细则。我们在一个涉及多司法管辖区推理与引用定位的真实法律问答任务上应用DeCE,评估不同大模型表现。结果表明,DeCE与专家评分的相关系数达0.78,显著高于传统指标(r=0.12)、点对点式大模型评分(r=0.35)及现代多维评估器(r=0.48)。同时,揭示出通用模型倾向召回,专用模型偏向精确。值得注意的是,仅11.95%的模型生成标准需专家修订,证明其高度可扩展性。DeCE为专家领域提供可解释、可操作的大模型评估方案。

原文摘要 · Abstract (English)

Evaluating long-form answers in high-stakes domains such as law or medicine remains a fundamental challenge. Standard metrics like BLEU and ROUGE fail to capture semantic correctness, and current LLM-based evaluators often reduce nuanced aspects of answer quality into a single undifferentiated score. We introduce DeCE, a decomposed LLM evaluation framework that separates precision (factual accuracy and relevance) and recall (coverage of required concepts), using instance-specific criteria automatically extracted from gold answer requirements. DeCE is model-agnostic and domain-general, requiring no predefined taxonomies or handcrafted rubrics. We instantiate DeCE to evaluate different LLMs on a real-world legal QA task involving multi-jurisdictional reasoning and citation grounding. DeCE achieves substantially stronger correlation with expert judgments ($r=0.78$), compared to traditional metrics ($r=0.12$), pointwise LLM scoring ($r=0.35$), and modern multidimensional evaluators ($r=0.48$). It also reveals interpretable trade-offs: generalist models favor recall, while specialized models favor precision. Importantly, only 11.95% of LLM-generated criteria required expert revision, underscoring DeCE's scalability. DeCE offers an interpretable and actionable LLM evaluation framework in expert domains.

大模型评估法律AI可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。