arXiv:2509.12405cs.CL2025-09被引 8

构建多语言医学问答评估基准,验证大模型评测更贴近专业判断。

MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering

  • 设计含2-4个专家答案的多语言医疗问答数据集
  • 大模型评测(如GPT-4)相关性显著高于传统指标
  • 适合医疗NLP评估、评测方法研究者使用

医学自然语言生成系统评估面临准确性、相关性和领域专长的严苛要求。传统自动评估指标(如BLEU、ROUGE、BERTScore)在开放式医学问答任务中难以区分高质量输出,因存在多个合理答案。本文提出MORQA(Medical Open-Response QA),一个包含英文和中文的多语言基准,涵盖三个医学图文问答数据集,每题有2-4个由医疗专家撰写的黄金标准答案,并附专家评分。我们对比传统指标与大语言模型(如GPT-4、Gemini)评估器,发现后者在与专家判断的相关性上显著更优。分析表明,大模型优势源于对语义细微差别的敏感性和对参考答案多样性的鲁棒性。本研究首次实现医学领域多语言、定性全面的生成评估分析,强调需采用人类对齐的评估方法。所有数据与标注将公开发布以支持后续研究。

原文摘要 · Abstract (English)

Evaluating natural language generation (NLG) systems in the medical domain presents unique challenges due to the critical demands for accuracy, relevance, and domain-specific expertise. Traditional automatic evaluation metrics, such as BLEU, ROUGE, and BERTScore, often fall short in distinguishing between high-quality outputs, especially given the open-ended nature of medical question answering (QA) tasks where multiple valid responses may exist. In this work, we introduce MORQA (Medical Open-Response QA), a new multilingual benchmark designed to assess the effectiveness of NLG evaluation metrics across three medical visual and text-based QA datasets in English and Chinese. Unlike prior resources, our datasets feature 2-4+ gold-standard answers authored by medical professionals, along with expert human ratings for three English and Chinese subsets. We benchmark both traditional metrics and large language model (LLM)-based evaluators, such as GPT-4 and Gemini, finding that LLM-based approaches significantly outperform traditional metrics in correlating with expert judgments. We further analyze factors driving this improvement, including LLMs' sensitivity to semantic nuances and robustness to variability among reference answers. Our results provide the first comprehensive, multilingual qualitative study of NLG evaluation in the medical domain, highlighting the need for human-aligned evaluation methods. All datasets and annotations will be publicly released to support future research.

医学NLP评估基准大模型评测多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。