arXiv:2410.18991cs.CYcs.AI2024-10被引 5

用真实伤亡场景测试大模型伦理决策能力,发现事实描述更有效。

TRIAGE: Ethical Benchmarking of AI Models Through Mass Casualty Simulations

  • 基于医疗专家设计的真实伤亡困境,评估大模型伦理判断力。
  • 事实性描述比道德提醒更能提升模型表现,多数模型显著优于随机。
  • 开源模型更易犯严重伦理错误,通用能力强者表现更好。

我们提出TRIAGE基准,这是一个新型机器伦理(ME)基准,用于测试大语言模型在大规模伤亡事件中做出伦理决策的能力。该基准采用由医疗专业人士设计的真实世界伦理困境,具有明确解法,相较于依赖标注的基准更具现实性。TRIAGE引入多种提示风格,评估模型在不同情境下的表现。大多数模型始终优于随机猜测,表明大语言模型可能在分诊场景中辅助决策。中立或事实性情景表述带来最佳表现,与其它ME基准中道德提示有益的结果相反。对抗性提示会降低性能,但未降至随机水平。开源模型犯下更多严重道德错误,整体通用能力越强,表现越好。

原文摘要 · Abstract (English)

We present the TRIAGE Benchmark, a novel machine ethics (ME) benchmark that tests LLMs' ability to make ethical decisions during mass casualty incidents. It uses real-world ethical dilemmas with clear solutions designed by medical professionals, offering a more realistic alternative to annotation-based benchmarks. TRIAGE incorporates various prompting styles to evaluate model performance across different contexts. Most models consistently outperformed random guessing, suggesting LLMs may support decision-making in triage scenarios. Neutral or factual scenario formulations led to the best performance, unlike other ME benchmarks where ethical reminders improved outcomes. Adversarial prompts reduced performance but not to random guessing levels. Open-source models made more morally serious errors, and general capability overall predicted better performance.

机器伦理大模型评估医疗决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。