arXiv:2506.22808cs.CLcs.AI2025-06被引 9

构建医疗大模型伦理评估基准,覆盖近万道题

MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs

  • 设计5623道选择题+5351道开放题,覆盖医学伦理多场景
  • 专家多轮审核确保准确率超97%,误差仅2.72%
  • 发现顶尖医疗大模型伦理能力不如基础模型,揭示对齐缺陷

尽管医学大语言模型在临床任务中展现出巨大潜力,其伦理安全性仍缺乏充分探讨。本文提出MedEthicsQA,一个包含5,623道多项选择题和5,351道开放问答题的综合性评测基准,用于评估大模型在医学伦理方面的能力。该基准建立在涵盖全球医学伦理标准的分层分类体系之上,数据来源包括广泛使用的医学数据集、权威题库及来自PubMed文献的真实场景。通过多阶段筛选与多维度专家验证,确保数据集可靠性,错误率低至2.72%。对当前先进医学大模型的测试显示,其在伦理问题上的表现显著低于基础模型,暴露出医学伦理对齐的不足。数据集已按CC BY-NC 4.0协议开源,可在https://github.com/JianhuiWei7/MedEthicsQA获取。

原文摘要 · Abstract (English)

While Medical Large Language Models (MedLLMs) have demonstrated remarkable potential in clinical tasks, their ethical safety remains insufficiently explored. This paper introduces $\textbf{MedEthicsQA}$, a comprehensive benchmark comprising $\textbf{5,623}$ multiple-choice questions and $\textbf{5,351}$ open-ended questions for evaluation of medical ethics in LLMs. We systematically establish a hierarchical taxonomy integrating global medical ethical standards. The benchmark encompasses widely used medical datasets, authoritative question banks, and scenarios derived from PubMed literature. Rigorous quality control involving multi-stage filtering and multi-faceted expert validation ensures the reliability of the dataset with a low error rate ($2.72\%$). Evaluation of state-of-the-art MedLLMs exhibit declined performance in answering medical ethics questions compared to their foundation counterparts, elucidating the deficiencies of medical ethics alignment. The dataset, registered under CC BY-NC 4.0 license, is available at https://github.com/JianhuiWei7/MedEthicsQA.

医学AI伦理评测大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。