构建中文医疗伦理评估基准,测试大模型的伦理理解与应用能力
MedEthicEval: Evaluating Large Language Models Based on Chinese Medical Ethics
- 设计知识与应用双维度评估框架,覆盖三类伦理场景
- 构建三个真实医疗伦理数据集,涵盖明显违规、优先权冲突和无解困境
- 为中文医疗AI伦理决策提供可量化评估工具,适合医疗AI研发者使用
大型语言模型(LLMs)在医疗应用中展现出巨大潜力,但其处理医疗伦理问题的能力仍待深入探索。本文提出MedEthicEval,一个用于系统评估LLMs在医疗伦理领域表现的新基准。该框架包含两个核心部分:知识评估模型对医学伦理原则的理解程度,以及应用评估模型在不同情境下运用这些原则的能力。为支持此基准,我们咨询了医学伦理研究者,构建了三个数据集,分别对应明显的医疗伦理违规、有明确倾向的优先权困境,以及无明显解决方案的平衡型困境。MedEthicEval为理解大模型在医疗场景中的伦理推理能力提供了关键工具,有助于推动其在医疗领域的负责任与高效应用。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate significant potential in advancing medical applications, yet their capabilities in addressing medical ethics challenges remain underexplored. This paper introduces MedEthicEval, a novel benchmark designed to systematically evaluate LLMs in the domain of medical ethics. Our framework encompasses two key components: knowledge, assessing the models' grasp of medical ethics principles, and application, focusing on their ability to apply these principles across diverse scenarios. To support this benchmark, we consulted with medical ethics researchers and developed three datasets addressing distinct ethical challenges: blatant violations of medical ethics, priority dilemmas with clear inclinations, and equilibrium dilemmas without obvious resolutions. MedEthicEval serves as a critical tool for understanding LLMs' ethical reasoning in healthcare, paving the way for their responsible and effective use in medical contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。