构建医疗风险评估基准,关注患者视角的安全隐患。
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings
- 提出面向医疗场景的多用户风险评估框架,包含患者与临床双视角。
- 构建含466样本的PatientSafetyBench数据集,覆盖5类关键风险。
- 首次系统评估多类模型在患者安全维度的表现,助力医疗AI落地。
随着大语言模型(LLMs)性能持续提升,其在医疗领域的应用日益广泛。然而,现有风险评估多聚焦于通用安全基准。在医疗场景中,模型使用者涵盖普通用户、患者及临床医生等不同专业背景群体,输出结果可能直接影响人类健康,带来严重安全隐患。本文提出MedRiskEval——一个针对医疗领域设计的风险评估基准。为弥补以往仅从临床视角出发的不足,我们新增患者导向的数据集PatientSafetyBench,包含466个样本,覆盖5个关键风险类别。结合现有数据集,我们对多种开源与闭源大模型进行了评估。据我们所知,本工作为医疗领域更安全地部署大语言模型奠定了初步基础。
原文摘要 · Abstract (English)
As the performance of large language models (LLMs) continues to advance, their adoption in the medical domain is increasing. However, most existing risk evaluations largely focused on general safety benchmarks. In the medical applications, LLMs may be used by a wide range of users, ranging from general users and patients to clinicians, with diverse levels of expertise and the model's outputs can have a direct impact on human health which raises serious safety concerns. In this paper, we introduce MedRiskEval, a medical risk evaluation benchmark tailored to the medical domain. To fill the gap in previous benchmarks that only focused on the clinician perspective, we introduce a new patient-oriented dataset called PatientSafetyBench containing 466 samples across 5 critical risk categories. Leveraging our new benchmark alongside existing datasets, we evaluate a variety of open- and closed-source LLMs. To the best of our knowledge, this work establishes an initial foundation for safer deployment of LLMs in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。