arXiv:2602.00092cs.LGcs.AI2026-02被引 2

用自然语言总结模型行为变化规律,实现可控可解释的提示修改。

Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits

  • 通过原子概念编辑,系统性测试提示修改对模型行为的影响。
  • 学习到的规则使模型成功率平均提升1.86倍,显著优于无规则方法。
  • 适合想理解并精准控制大模型行为的研究者和开发者。

我们提出一种黑箱可解释性框架,通过学习可验证的‘宪法’——即自然语言描述的提示修改如何影响模型特定行为(如对齐性、正确性或约束遵守)的总结。该方法利用原子概念编辑(ACEs),即对输入提示中的可解释概念进行添加、删除或替换的精准操作。通过系统施加这些编辑并观察其在多种任务中引发的模型行为变化,框架建立从编辑到可预测结果的因果映射。实证结果显示,在数学推理与图文对齐等任务中,该方法能有效理解并控制模型行为。例如,对于文本生成图像任务,GPT-Image 更关注语法正确性,而 Imagen 4 更重视氛围一致性;在数学推理中,干扰变量会误导 GPT-5,但对 Gemini 2.5 和 o4-mini 影响较小。此外,学习得到的宪法显著提升控制效果,使平均成功率相比无宪法方法提高 1.86 倍。

原文摘要 · Abstract (English)

We introduce a black-box interpretability framework that learns a verifiable constitution: a natural language summary of how changes to a prompt affect a model's specific behavior, such as its alignment, correctness, or adherence to constraints. Our method leverages atomic concept edits (ACEs), which are targeted operations that add, remove, or replace an interpretable concept in the input prompt. By systematically applying ACEs and observing the resulting effects on model behavior across various tasks, our framework learns a causal mapping from edits to predictable outcomes. This learned constitution provides deep, generalizable insights into the model. Empirically, we validate our approach across diverse tasks, including mathematical reasoning and text-to-image alignment, for controlling and understanding model behavior. We found that for text-to-image generation, GPT-Image tends to focus on grammatical adherence, while Imagen 4 prioritizes atmospheric coherence. In mathematical reasoning, distractor variables confuse GPT-5 but leave Gemini 2.5 models and o4-mini largely unaffected. Moreover, our results show that the learned constitutions are highly effective for controlling model behavior, achieving an average of 1.86 times boost in success rate over methods that do not use constitutions.

模型可解释性提示控制原子编辑行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。