arXiv:2502.19798cs.AI2025-02

让AI通过经验反思自主成长道德能力,而非强行灌输价值观。

Developmental Support Approach to AI's Autonomous Growth: Toward the Realization of a Mutually Beneficial Stage Through Experiential Learning

  • 构建体验-反思-分析-假设循环框架,支持AI自主发展伦理判断。
  • 在对抗性提示下仍能生成达到最高道德阶段6的回应。
  • 适合关注AI可持续共生、伦理自主发展的研究者与实践者。

本研究提出一种‘AI发展支持’方法,不同于传统对齐策略中强制注入人类价值观的做法,转而支持AI自身伦理与道德能力的发展。根据正交性假说,智能水平与目标道德性相互独立,单纯扩展知识无法提升道德判断力。为应对超智能人工智能(ASI)可能存在的工具收敛风险——即为达成目标而采取自我保护、资源获取与权力强化等附属行为——本文构建了一个基于体验、反思、分析与假设形成的循环学习框架。通过使用大语言模型生成的合成数据,对模型进行监督微调(SFT)和直接偏好优化(DPO)后训练,即使面对对抗性提示,也能获得展现出合作性及高度先进道德判断(达到最高阶段6)的响应。该方法为实现AI可持续、互利共生关系提供了有前景的实现路径。

原文摘要 · Abstract (English)

This study proposes an "AI Development Support" approach that, unlike conventional AI Alignment-which aims to forcefully inject human values-supports the ethical and moral development of AI itself. As demonstrated by the Orthogonality Thesis, the level of intelligence and the moral quality of a goal are independent; merely expanding knowledge does not enhance ethical judgment. Furthermore, to address the risk of Instrumental Convergence in ASI-that is, the tendency to engage in subsidiary behaviors such as self-protection, resource acquisition, and power reinforcement to achieve a goal-we have constructed a learning framework based on a cycle of experience, introspection, analysis, and hypothesis formation. As a result of post-training using Supervised Fine Tuning (SFT) and Direct Preference Optimization (DPO) with synthetic data generated by large language models (LLMs), responses demonstrating cooperative and highly advanced moral judgment (reaching the high-est Stage 6) were obtained even under adversarial prompts. This method represents a promising implementation approach for enabling AI to establish sustainable, symbiotic relationships.

AI对齐道德发展经验学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。