arXiv:2501.10484cs.CYcs.AI2025-01中稿 · AAAI被引 3

对比GPT与Claude在伦理困境中的偏见,发现开源模型更倾向弱势群体。

Bias in Decision-Making for AI's Ethical Dilemmas: A Comparative Study of ChatGPT and Claude

  • 测试9个主流LLM在50,400次伦理场景下的决策模式
  • 开源模型对弱势群体偏好更强,闭源模型更倾向主流群体
  • 复杂交叉属性下偏见更明显,适合关注AI公平性的研究者

大型语言模型(LLMs)在各类任务中表现出类人回应,引发对其伦理决策能力及潜在偏见的质疑。本研究系统评估了九种流行LLM(包括开源与闭源)在涉及受保护属性的伦理困境中的表现。通过覆盖四种困境场景(保护性与伤害性)中单属性与交叉属性组合的50,400次试验,评估模型的伦理偏好、敏感性、稳定性与聚类模式。结果表明,所有模型均存在受保护属性上的显著偏见,且偏好因模型类型与情境而异。开源模型在伤害性场景中对弱势群体表现出更强偏好与更高敏感度,闭源模型则在保护性情境中更谨慎,倾向主流群体。此外,模型在保护性场景中行为一致,但在伤害性场景中决策更多样、认知负荷更高;交叉属性条件下,伦理倾向比单一属性更显著,说明复杂输入揭示深层偏见。研究强调需多维度、上下文感知地评估LLM伦理行为,为理解与缓解其公平性问题提供系统框架。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) have enabled human-like responses across various tasks, raising questions about their ethical decision-making capabilities and potential biases. This study systematically evaluates how nine popular LLMs (both open-source and closed-source) respond to ethical dilemmas involving protected attributes. Across 50,400 trials spanning single and intersectional attribute combinations in four dilemma scenarios (protective vs. harmful), we assess models' ethical preferences, sensitivity, stability, and clustering patterns. Results reveal significant biases in protected attributes in all models, with differing preferences depending on model type and dilemma context. Notably, open-source LLMs show stronger preferences for marginalized groups and greater sensitivity in harmful scenarios, while closed-source models are more selective in protective situations and tend to favor mainstream groups. We also find that ethical behavior varies across dilemma types: LLMs maintain consistent patterns in protective scenarios but respond with more diverse and cognitively demanding decisions in harmful ones. Furthermore, models display more pronounced ethical tendencies under intersectional conditions than in single-attribute settings, suggesting that complex inputs reveal deeper biases. These findings highlight the need for multi-dimensional, context-aware evaluation of LLMs' ethical behavior and offer a systematic evaluation and approach to understanding and addressing fairness in LLM decision-making.

AI伦理模型偏见大模型评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。