用七种人类道德视角协调AI决策冲突,避免单一标准压制多元价值。
PRISM: Perspective Reasoning for Integrated Synthesis and Mediation as a Multi-Perspective Framework for AI Alignment
- 基于认知科学构建七类道德视角,系统捕捉人类多元价值。
- 通过帕累托式优化平衡冲突,不简化为单一指标。
- 适合处理公共政策、教育等高伦理复杂度场景的AI对齐问题。
本文提出多视角框架PRISM,用于解决人工智能对齐中的持久挑战,如人类价值观冲突与规范规避问题。该框架基于认知科学与道德心理学,将道德关切归纳为七个‘基础世界观’,涵盖从生存本能到高层次整合的多个维度。通过受帕累托启发的优化机制,在不将多元目标降维为单一指标的前提下,协调竞争性优先级。在可靠上下文验证的前提下,框架遵循结构化流程:获取各视角特异性响应,合成平衡结果,并以透明、迭代方式调解剩余冲突。结合认知科学、道德心理学与神经科学的多层次道德认知理论,PRISM阐明了不同道德驱动力的交互机制,系统记录并调解伦理权衡。通过工作原型实证展示其在公共卫生政策、职场自动化、教育等经典对齐问题中的有效性。通过锚定人类视角,限制可能偏离人类立场的解释跳跃。简要展望未来方向,包括现实部署与形式验证,核心聚焦于多视角融合与冲突调解。
原文摘要 · Abstract (English)
In this work, we propose Perspective Reasoning for Integrated Synthesis and Mediation (PRISM), a multiple-perspective framework for addressing persistent challenges in AI alignment such as conflicting human values and specification gaming. Grounded in cognitive science and moral psychology, PRISM organizes moral concerns into seven "basis worldviews", each hypothesized to capture a distinct dimension of human moral cognition, ranging from survival-focused reflexes through higher-order integrative perspectives. It then applies a Pareto-inspired optimization scheme to reconcile competing priorities without reducing them to a single metric. Under the assumption of reliable context validation for robust use, the framework follows a structured workflow that elicits viewpoint-specific responses, synthesizes them into a balanced outcome, and mediates remaining conflicts in a transparent and iterative manner. By referencing layered approaches to moral cognition from cognitive science, moral psychology, and neuroscience, PRISM clarifies how different moral drives interact and systematically documents and mediates ethical tradeoffs. We illustrate its efficacy through real outputs produced by a working prototype, applying PRISM to classic alignment problems in domains such as public health policy, workplace automation, and education. By anchoring AI deliberation in these human vantage points, PRISM aims to bound interpretive leaps that might otherwise drift into non-human or machine-centric territory. We briefly outline future directions, including real-world deployments and formal verifications, while maintaining the core focus on multi-perspective synthesis and conflict mediation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。