用多学科协作让AI做道德判断更可靠透明
CogniAlign: Survivability-Grounded Multi-Agent Moral Reasoning for Safe and Transparent AI
- 让不同领域科学家角色辩论,以生存性为道德基础
- 在60多个道德问题上得分比GPT-4o高12.2至31.2分
- 适合关注AI安全与可解释性的研究者和开发者
当前人工智能对齐人类价值观的挑战源于道德原则的抽象性与冲突性,以及现有方法的不透明。本文提出基于自然主义道德实在论的多智能体协商框架CogniAlign,将道德推理建立在个体与集体层面的生存性基础上,并通过神经科学、心理学、社会学和进化生物学等专业领域智能体的结构化讨论实现。每个智能体提供论证与反驳,由仲裁者整合生成透明且有实证依据的判断。作为概念验证,我们在经典与新设道德问题上评估CogniAlign,结合三位专家使用五部分伦理审计框架与GPT-4o对比。结果表明,在超过六十个道德问题上,CogniAlign平均在分析质量上高出12.2分,决断力提升31.2分,解释深度增加15分。例如在海因茨困境中,其总分为79,显著优于GPT-4o的65.8。该方法展示了可审计的AI对齐路径可行性,尽管仍存在若干挑战。
原文摘要 · Abstract (English)
The challenge of aligning artificial intelligence (AI) with human values persists due to the abstract and often conflicting nature of moral principles and the opacity of existing approaches. This paper introduces CogniAlign, a multi-agent deliberation framework based on naturalistic moral realism, that grounds moral reasoning in survivability, defined across individual and collective dimensions, and operationalizes it through structured deliberations among discipline-specific scientist agents. Each agent, representing neuroscience, psychology, sociology, and evolutionary biology, provides arguments and rebuttals that are synthesized by an arbiter into transparent and empirically anchored judgments. As a proof-of-concept study, we evaluate CogniAlign on classic and novel moral questions and compare its outputs against GPT-4o using a five-part ethical audit framework with the help of three experts. Results show that CogniAlign consistently outperforms the baseline across more than sixty moral questions, with average performance gains of 12.2 points in analytic quality, 31.2 points in decisiveness, and 15 points in depth of explanation. In the Heinz dilemma, for example, CogniAlign achieved an overall score of 79 compared to GPT-4o's 65.8, demonstrating a decisive advantage in handling moral reasoning. Through transparent and structured reasoning, CogniAlign demonstrates the feasibility of an auditable approach to AI alignment, though certain challenges still remain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。