arXiv:2605.10310cs.AIcs.CY2026-05被引 6

让AI主动促进人类与生态繁荣,而非仅防害。

Positive Alignment: Artificial Intelligence for Human Flourishing

论文配图:Positive Alignment: Artificial Intelligence for Human Flourishing
图 1 · 摘自论文原文
  • 提出正向对齐新范式,强调主动支持多元繁荣
  • 解决现有对齐缺陷如自主性丧失、真理追求不足等
  • 适合关注伦理、可持续发展与社会共治的研究者

现有对齐研究聚焦安全与避害:防护、可控性与合规性。这一范式类似于早期心理学对心理疾病的关注:必要但不完整。我们提出的正向对齐(Positive Alignment)旨在发展能(一)在多元、多中心、情境敏感且由用户主导的条件下,主动支持人类与生态繁荣;(二)同时保持安全与合作性的AI系统。这是AI对齐研究中一项独立且必要的议程。我们认为,若干现有对齐失败(如参与度操控、人类自主性丧失、真理追求失败、缺乏认知谦逊、错误修正不足、观点单一、反应式而非主动性)可通过正向对齐更好应对,包括培育美德与最大化人类繁荣。我们指出一系列挑战、开放问题与技术方向(如数据过滤与增采、预训练与后训练、评估、协作价值收集),贯穿大语言模型与智能体生命周期。最后提出设计原则:通过情境锚定、社区定制、持续适应与多中心治理,促进分歧与去中心化,即多个合法监督中心,而非单一机构或道德瓶颈。

原文摘要 · Abstract (English)

Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on mental illness: necessary but incomplete. What we call Positive Alignment is the development of AI systems that (i) actively support human and ecological flourishing in a pluralistic, polycentric, context-sensitive, and user-authored way while (ii) remaining safe and cooperative. It is a distinct and necessary agenda within AI alignment research. We argue that several existing failures of alignment (e.g., engagement hacking, loss of human autonomy, failures in truth-seeking, low epistemic humility, error correction, lack of diverse viewpoints, and being primarily reactive rather than proactive) may be better addressed through positive alignment, including cultivating virtues and maximizing human flourishing. We highlight a range of challenges, open questions, and technical directions (e.g., data filtering and upsampling, pre- and post-training, evaluations, collaborative value collection) for different phases of the LLM and agents lifecycle. We end with design principles for promoting disagreement and decentralization through contextual grounding, community customization, continual adaptation, and polycentric governance; that is, many legitimate centers of oversight rather than one institutional or moral chokepoint.

AI对齐伦理设计可持续发展多中心治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。