提出让AI适度偏离人类意图,以形成制衡生态来应对对齐难题。
Neurodivergent Influenceability as a Contingent Solution to the AI Alignment Problem
- 认为完全对齐不可能,主张利用不可避免的偏差构建竞争生态
- 实验表明开放模型更多样,闭源模型更易控制且可对抗对手
- 人与AI干预效果不同,提示需多策略协同应对风险
AI对齐问题关乎人工智能(包括通用智能与超智能)是否遵循人类价值观,随着技术从专用AI向通用智能(AGI)和超智能(ASI)演进,控制与生存风险日益加剧。本文探讨了拥抱不可避免的对齐偏差,作为推动多元竞争主体共存、引导系统朝更符合人类利益方向发展的可行策略。核心论点为:由于图灵完备系统存在数学上无法实现完全对齐,该特性将继承至AGI与ASI系统。为此,我们设计基于扰动与干预分析的‘观念转变攻击测试’,研究人类与智能体如何通过合作或竞争改变友好或敌意智能体的行为。结果表明,开放模型更具多样性;专有模型虽能有效施加约束(有正负两面后果),但封闭系统更易调控,并可用于对抗专有系统。此外,人类与智能体干预效果各异,提示需采用差异化策略。
原文摘要 · Abstract (English)
The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General Intelligence (AGI) and Superintelligence, fears about control and existential risk have escalated. Here, we investigate whether embracing inevitable AI misalignment can be a contingent strategy to foster a dynamic ecosystem of competing agents as a viable path to steer them in more human-aligned trends and mitigate risks. We explore how misalignment may serve and should be promoted as a counterbalancing mechanism to team up with whichever agents are most aligned to human interests, ensuring that no single system dominates destructively. The main premise of our contribution is that misalignment is inevitable because full AI-human alignment is a mathematical impossibility from Turing-complete systems, which we also offer as a proof in this contribution, a feature then inherited to AGI and ASI systems. We introduce a change-of-opinion attack test based on perturbation and intervention analysis to study how humans and agents may change or neutralise friendly and unfriendly AIs through cooperation and competition. We show that open models are more diverse and that most likely guardrails implemented in proprietary models are successful at controlling some of the agents' range of behaviour with positive and negative consequences while closed systems are more steerable and can also be used against proprietary AI systems. We also show that human and AI intervention has different effects hence suggesting multiple strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。