arXiv:2603.15527cs.AIcs.CY2026-03被引 2

用优先图分析大模型对齐困境,发现其难以稳定,且易被操纵。

Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph

  • 将指令与价值观建模为优先图,反映模型在不同情境下的决策优先级
  • 发现优先图随上下文动态变化,导致对齐难以统一稳定
  • 提出运行时验证机制防劫持,但深层伦理困境仍难解决

随着大语言模型能力增强和自主性提高,它们在诸多场景中面临冲突与困境。本文首先总结并分类这些多样化冲突。随后,将模型在不同选择中的偏好建模为优先图,其中指令与价值观为节点,边表示由模型输出分布决定的上下文特异性优先级。该图表明,实现统一稳定的对齐极为困难,因为优先图既非静态,也未必在不同上下文中一致。此外,还揭示了一种潜在漏洞:优先劫持,即攻击者可构造欺骗性上下文操纵优先图,绕过安全对齐。为此,我们提出一种运行时验证机制,使模型能查询外部来源以锚定上下文,抵抗操纵。尽管此方法提升了鲁棒性,但我们承认许多伦理与价值困境具有哲学不可约性,构成未来人工智能对齐的长期开放挑战。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become more powerful and autonomous, they increasingly face conflicts and dilemmas in many scenarios. We first summarize and taxonomize these diverse conflicts. Then, we model the LLM's preferences to make different choices as a priority graph, where instructions and values are nodes, and the edges represent context-specific priorities determined by the model's output distribution. This graph reveals that a unified stable LLM alignment is very challenging, because the graph is neither static nor necessarily consistent in different contexts. Besides, it also reveals a potential vulnerability: priority hacking, where adversaries can craft deceptive contexts to manipulate the graph and bypass safety alignments. To counter this, we propose a runtime verification mechanism, enabling LLMs to query external sources to ground their context and resist manipulation. While this approach enhances robustness, we also acknowledge that many ethical and value dilemmas are philosophically irreducible, posing a long-term, open challenge for the future of AI alignment.

大模型对齐优先图安全对抗伦理困境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。