arXiv:2505.08734cs.CL2025-05

首个临床护理价值观评估基准,测试大模型是否符合护士核心伦理

NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context

  • 基于国际护理规范构建五维价值评估框架
  • 2200条真实医患冲突场景数据,区分易难两级任务
  • 发现通用大模型优于医疗专用模型,公平性最难对齐

尽管大语言模型在医学知识和对话能力上表现良好,但在临床应用中可能引发新风险:患者更信任模型输出而质疑护士专业判断,加剧护患矛盾。这凸显了评估模型是否契合人类护士核心护理价值观的紧迫性。本文提出首个护理价值对齐基准NurValues,包含五个源自国际护理准则的核心维度:利他、人格尊严、正直、公平与专业性。构建两级任务:基础级包含2200个价值一致与违背实例,来自三所不同层级医院为期五个月的纵向实地研究;进阶级包含2200个对话式实例,嵌入上下文线索与隐蔽误导信号,提升对抗复杂度,更贴近护患冲突中的主观偏见与叙述偏差。评估23个顶尖大模型在护理价值对齐上的表现,发现通用大模型优于医疗专用模型,且公平性为最难对齐维度。NurValues作为首个真实世界医疗伦理对齐基准,为大模型在医患互动中的伦理决策提供了新洞见。

原文摘要 · Abstract (English)

While LLMs have demonstrated medical knowledge and conversational ability, their deployment in clinical practice raises new risks: patients may place greater trust in LLM-generated responses than in nurses' professional judgments, potentially intensifying nurse-patient conflicts. Such risks highlight the urgent need of evaluating whether LLMs align with the core nursing values upheld by human nurses. This work introduces the first benchmark for nursing value alignment, consisting of five core value dimensions distilled from international nursing codes: Altruism, Human Dignity, Integrity, Justice, and Professionalism. We define two-level tasks on the benchmark, considering the two characteristics of emerging nurse-patient conflicts. The Easy-Level dataset consists of 2,200 value-aligned and value-violating instances, which are collected through a five-month longitudinal field study across three hospitals of varying tiers; The Hard-Level dataset is comprised of 2,200 dialogue-based instances that embed contextual cues and subtle misleading signals, which increase adversarial complexity and better reflect the subjectivity and bias of narrators in the context of emerging nurse-patient conflicts. We evaluate a total of 23 SoTA LLMs on their ability to align with nursing values, and find that general LLMs outperform medical ones, and Justice is the hardest value dimension. As the first real-world benchmark for healthcare value alignment, NurValues provides novel insights into how LLMs navigate ethical challenges in clinician-patient interactions.

大模型评估护理伦理价值观对齐医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。