用五维模型评估大模型的伦理推理能力,发现解释力差异大。
Auditing the Ethical Logic of Generative AI Models
- 构建五维审计框架:分析质量、伦理广度、解释深度、一致性与决断力。
- 七款主流大模型在伦理判断上趋同,但解释严谨性与道德优先级差异显著。
- 思维链提示和优化推理的模型显著提升审计表现,适合伦理评测研究者。
随着生成式AI日益融入高风险领域,对其伦理推理能力进行可靠评估变得愈发重要。本文提出一个五维审计模型,用于评估主流大语言模型(LLMs)的伦理逻辑,涵盖分析质量、伦理考量广度、解释深度、一致性和决断力。基于应用伦理学与高阶思维传统,设计多套提示测试,包含新颖的伦理困境,以探测模型在多样化情境下的推理表现。对七款主要大模型的基准测试显示,尽管模型在伦理决策上普遍趋同,但在解释严谨性与道德优先排序方面存在显著差异。使用思维链(Chain-of-Thought)提示及推理优化模型可显著提升审计指标表现。本研究提出了可扩展的AI伦理评估方法,并揭示了AI在复杂决策中辅助人类道德判断的潜力。
原文摘要 · Abstract (English)
As generative AI models become increasingly integrated into high-stakes domains, the need for robust methods to evaluate their ethical reasoning becomes increasingly important. This paper introduces a five-dimensional audit model -- assessing Analytic Quality, Breadth of Ethical Considerations, Depth of Explanation, Consistency, and Decisiveness -- to evaluate the ethical logic of leading large language models (LLMs). Drawing on traditions from applied ethics and higher-order thinking, we present a multi-battery prompt approach, including novel ethical dilemmas, to probe the models' reasoning across diverse contexts. We benchmark seven major LLMs finding that while models generally converge on ethical decisions, they vary in explanatory rigor and moral prioritization. Chain-of-Thought prompting and reasoning-optimized models significantly enhance performance on our audit metrics. This study introduces a scalable methodology for ethical benchmarking of AI systems and highlights the potential for AI to complement human moral reasoning in complex decision-making contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。