提出四维无参考评估框架,精准衡量大模型在专业场景下的回答质量。
SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation
- 构建四维评估体系:具体性、抗改写鲁棒性、相关性与上下文利用度。
- 在1412个专业问答对上验证,单一指标无法全面反映答案质量。
- 适用于灾害响应、基建规划等高风险领域的模型评估,适合研究者与从业者使用。
大型语言模型(LLMs)正广泛应用于自然灾害应对与基础设施规划等高风险领域,其回答需包含细粒度的决策关键信息。然而,现有检索增强生成(RAG)与开放式问答评估多依赖表面相似性、事实一致性或语义相关性,难以衡量响应是否提供决策必需的具体信息。为此,我们提出一种多维度、无参考的评估框架,从具体性、对改写与语义扰动的鲁棒性、答案相关性及上下文利用度四个互补维度评估输出。我们构建了一个涵盖40种专业角色和7类自然灾害的1,412个领域特定问答对的数据集以支持系统性评估。通过人工评估验证标注者间一致性和模型输出与人类判断的匹配度,揭示了开放式、领域特定评估的内在主观性。结果表明,单一指标无法充分捕捉答案质量,强调在高风险应用中部署大模型时需采用结构化多指标评估框架。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to support question answering and decision-making in high-stakes, domain-specific settings such as natural hazard response and infrastructure planning, where effective answers must convey fine-grained, decision-critical details. However, existing evaluation frameworks for retrieval-augmented generation (RAG) and open-ended question answering primarily rely on surface-level similarity, factual consistency, or semantic relevance, and often fail to assess whether responses provide the specific information required for domain-sensitive decisions. To address this gap, we propose a multi-dimensional, reference-free evaluation framework that assesses LLM outputs along four complementary dimensions: specificity, robustness to paraphrasing and semantic perturbations, answer relevance, and context utilization. We introduce a curated dataset of 1,412 domain-specific question-answer pairs spanning 40 professional roles and seven natural hazard types to support systematic evaluation. We further conduct human evaluation to assess inter-annotator agreement and alignment between model outputs and human judgments, which highlights the inherent subjectivity of open-ended, domain-specific evaluation. Our results show that no single metric sufficiently captures answer quality in isolation and demonstrate the need for structured, multi-metric evaluation frameworks when deploying LLMs in high-stakes applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。