arXiv:2602.18823cs.CL2026-02Conference of the …

让大模型评估更可靠,专为医疗等敏感领域设计。

EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation

  • 提供交互式引导和自动元评估工具,选对评测方法
  • 用扰动数据验证评测结果可靠性,避免偏差
  • 支持多种模型和策略,适合医疗等专业场景

大语言模型的稳健评估对识别有效系统配置、降低敏感领域部署风险至关重要。传统统计指标难以适用于开放式生成任务,导致越来越多依赖基于大模型的评估方法。这类方法虽灵活,但依赖特定模型、提示词、参数和策略,易因配置不当引入偏差。本文提出 EvalSense,一个可扩展的领域专用大模型评估框架。它提供对多种模型提供商和评估策略的开箱即用支持,并通过两个核心组件辅助用户:(1)交互式评估方法选择引导;(2)基于扰动数据的自动化元评估工具,用于检验不同评估方法的可靠性。我们在临床病历生成任务中开展案例研究,使用公开的医生-患者对话数据集验证其有效性。所有代码、文档和资源已开源,可在 https://github.com/nhsengland/evalsense 获取。

原文摘要 · Abstract (English)

Robust and comprehensive evaluation of large language models (LLMs) is essential for identifying effective LLM system configurations and mitigating risks associated with deploying LLMs in sensitive domains. However, traditional statistical metrics are poorly suited to open-ended generation tasks, leading to growing reliance on LLM-based evaluation methods. These methods, while often more flexible, introduce additional complexity: they depend on carefully chosen models, prompts, parameters, and evaluation strategies, making the evaluation process prone to misconfiguration and bias. In this work, we present EvalSense, a flexible, extensible framework for constructing domain-specific evaluation suites for LLMs. EvalSense provides out-of-the-box support for a broad range of model providers and evaluation strategies, and assists users in selecting and deploying suitable evaluation methods for their specific use-cases. This is achieved through two unique components: (1) an interactive guide aiding users in evaluation method selection and (2) automated meta-evaluation tools that assess the reliability of different evaluation approaches using perturbed data. We demonstrate the effectiveness of EvalSense in a case study involving the generation of clinical notes from unstructured doctor-patient dialogues, using a popular open dataset. All code, documentation, and assets associated with EvalSense are open-source and publicly available at https://github.com/nhsengland/evalsense.

大模型评估医疗AI元评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。