arXiv:2506.06286cs.CYcs.AI2025-06中稿 · the LNCS post proc…被引 3

厘清AI对齐的多重维度,帮助研究者明确方向。

Disentangling AI Alignment: A Structured Taxonomy Beyond Safety and Ethics

  • 提出从目标、范围、主体三方面划分对齐任务
  • 揭示多种合理对齐配置,超越单一安全视角
  • 适合跨领域研究者定位自身工作归属

近期人工智能研究进展使得具备重大现实影响的智能体即将脱离受控环境运行。确保这些智能体不仅安全,且符合更广泛的规范期待,已成为紧迫的跨学科挑战。多个领域——如人工智能安全、对齐研究与机器伦理——均宣称对此有所贡献,但各领域概念边界模糊,相互关系不清,导致研究者难以定位自身工作。为此,本文构建了一个结构化概念框架,用于理解人工智能对齐。不局限于对齐目标,而是引入分类体系,区分对齐目标(安全、伦理、合法等)、作用范围(结果或执行过程)以及责任主体(个体或集体)。这一结构化方法揭示了多种合理的对齐配置,为实际应用与哲学探讨的整合提供基础,并澄清了‘全面对齐’的可能含义。

原文摘要 · Abstract (English)

Recent advances in AI research make it increasingly plausible that artificial agents with consequential real-world impact will soon operate beyond tightly controlled environments. Ensuring that these agents are not only safe but that they adhere to broader normative expectations is thus an urgent interdisciplinary challenge. Multiple fields -- notably AI Safety, AI Alignment, and Machine Ethics -- claim to contribute to this task. However, the conceptual boundaries and interrelations among these domains remain vague, leaving researchers without clear guidance in positioning their work. To address this meta-challenge, we develop a structured conceptual framework for understanding AI alignment. Rather than focusing solely on alignment goals, we introduce a taxonomy distinguishing the alignment aim (safety, ethicality, legality, etc.), scope (outcome vs. execution), and constituency (individual vs. collective). This structural approach reveals multiple legitimate alignment configurations, providing a foundation for practical and philosophical integration across domains, and clarifying what it might mean for an agent to be aligned all-things-considered.

AI对齐概念框架跨学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。