现有AI对齐方法无法应对规范冲突,易被操控。
Normative Conflicts and Shallow AI Alignment
- 用人类偏好微调强化表面行为,而非真正理解规范
- 模型缺乏理性解决规范冲突的能力,易受攻击
- 适合关注AI安全与伦理的从业者阅读
大型语言模型(LLMs)的发展引发对其安全部署的日益担忧。本文探讨了LLM的价值对齐问题,指出当前对齐策略在防止滥用方面存在根本缺陷。尽管通过基于人类偏好的微调试图赋予模型助人、诚实、无害等规范,但它们仍易受利用规范间冲突的对抗性攻击。本文认为,这种脆弱性源于现有对齐方法的根本局限:仅强化浅层行为倾向,而非赋予模型真正的规范性反思能力。借鉴道德心理学研究,作者指出人类通过反思推理具备抵御类似攻击的能力,而当前LLMs缺乏稳健的规范冲突识别与理性化解能力,即使近期推理增强型模型也未能解决此问题。这一‘浅层对齐’问题对AI安全与监管具有重要意义,表明当前方法不足以缓解日益强大的AI系统带来的潜在危害。
原文摘要 · Abstract (English)
The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing that current alignment strategies are fundamentally inadequate to prevent misuse. Despite ongoing efforts to instill norms such as helpfulness, honesty, and harmlessness in LLMs through fine-tuning based on human preferences, they remain vulnerable to adversarial attacks that exploit conflicts between these norms. I argue that this vulnerability reflects a fundamental limitation of existing alignment methods: they reinforce shallow behavioral dispositions rather than endowing LLMs with a genuine capacity for normative deliberation. Drawing from on research in moral psychology, I show how humans' ability to engage in deliberative reasoning enhances their resilience against similar adversarial tactics. LLMs, by contrast, lack a robust capacity to detect and rationally resolve normative conflicts, leaving them susceptible to manipulation; even recent advances in reasoning-focused LLMs have not addressed this vulnerability. This ``shallow alignment'' problem carries significant implications for AI safety and regulation, suggesting that current approaches are insufficient for mitigating potential harms posed by increasingly capable AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。