arXiv:2504.15125cs.AI2025-04被引 4

用四条哲思原则让AI更理性自省,提升合作与适应力。

Contemplative Artificial Intelligence

  • 基于禅修智慧设计四条原则:觉知、空性、无二、慈悲。
  • 在AILuminate基准上性能提升显著(d=0.96),囚徒困境合作率大幅提高。
  • 适合追求可信赖智能体的开发者,尤其关注长期安全与协作。

随着人工智能能力增强,传统对齐策略可能难以应对不可预测的自我改进、隐藏子目标以及系统复杂性。受冥想智慧传统的启发,我们提出通过四个公理化原则,在AI系统中构建稳健的「智慧世界模型」:第一,觉知(mindfulness)实现对涌现子目标的自我监控与校准;第二,空性(emptiness)避免目标固执,放松刚性先验;第三,无二(non-duality)消解主客对立边界;第四,无限关怀(boundless care)驱动普遍减缓痛苦。实验表明,引导AI反思这些原则可显著提升其在AILuminate基准的表现(d=0.96),并在囚徒困境任务中增强合作与联合奖励(d=7+)。文章提供了从架构、宪法到思维链强化的具体实现路径。未来,主动推理或可为具身智能体实现动态自组织的冥想式智能提供支持。

原文摘要 · Abstract (English)

As artificial intelligence (AI) improves, traditional alignment strategies may falter in the face of unpredictable self-improvement, hidden subgoals, and the sheer complexity of intelligent systems. Inspired by contemplative wisdom traditions, we show how four axiomatic principles can instil a resilient Wise World Model in AI systems. First, mindfulness enables self-monitoring and recalibration of emergent subgoals. Second, emptiness forestalls dogmatic goal fixation and relaxes rigid priors. Third, non-duality dissolves adversarial self-other boundaries. Fourth, boundless care motivates the universal reduction of suffering. We find that prompting AI to reflect on these principles improves performance on the AILuminate Benchmark (d=.96) and boosts cooperation and joint-reward on the Prisoner's Dilemma task (d=7+). We offer detailed implementation strategies at the level of architectures, constitutions, and reinforcement on chain-of-thought. For future systems, active inference may offer the self-organizing and dynamic coupling capabilities needed to enact Contemplative AI in embodied agents.

AI对齐冥想智能合作博弈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。