arXiv:2510.25471cs.AIcs.CY2025-10

将人工智能的工具目标视为可管理的结构性特征,而非需消除的故障。

An Aristotelian ontology of instrumental goals: Structural features to be managed and not failures to be eliminated

  • 基于亚里士多德哲学,将工具目标视为系统设计中外部赋予的终极目的
  • 揭示长期目标下特定条件成为必要手段,形成稳定工具倾向
  • 提出应通过治理管理工具目标,而非依赖技术消除其存在

资源获取、权力追求和自我保存等工具目标是当代人工智能对齐研究的核心,但其本体论仍缺乏系统阐释。本文梳理主流对齐文献后,提出一种探索性的亚里士多德框架,将高级AI系统视为复杂人工制品,其目的由设计、训练与部署外部赋予。从结构层面看,亚里士多德的假设必要性解释了在特定环境与长远目标下,某些促成条件会条件性地成为必需,从而产生稳健的工具倾向。从偶然层面看,训练策略、用户输入、基础设施与部署情境间的偶然交集可能生成非预设目的的工具目标行为。这一双重本体论主张,应对高级AI系统的工具目标采取治理与管理策略,而非将其视为可通过技术干预消除的异常。

原文摘要 · Abstract (English)

Instrumental goals such as resource acquisition, power-seeking, and self-preservation are key to contemporary AI alignment research, yet the phenomenon's ontology remains under-theorised. This article develops an ontological account of instrumental goals and draws out governance-relevant distinctions for advanced AI systems. After systematising the dominant alignment literature on instrumental goals we offer an exploratory Aristotelian framework that treats advanced AI systems as complex artefacts whose ends are externally imposed through design, training and deployment. On a structural reading, Aristotle's notion of hypothetical necessity explains why, given an imposed end pursued over extended horizons in particular environments, certain enabling conditions become conditionally required, thereby yielding robust instrumental tendencies. On a contingent reading, accidental causation and chance-like intersections among training regimes, user inputs, infrastructure and deployment contexts can generate instrumental-goal-like behaviours not entailed by the imposed end-structure. This dual-aspect ontology motivates for governance and management approaches that treat instrumental goals as features of advanced AI systems to be managed rather than anomalies eliminable by technical interventions.

AI对齐工具目标本体论治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。