找到让大模型生成特定输出的最小神经元组合,实现精准可控。
WASD: Locating Critical Neurons as Sufficient Conditions for Explaining and Controlling LLM Behavior
- 通过激活神经元条件定位关键行为原因。
- 在SST-2和CounterFact上比传统方法更稳定准确。
- 适合需要精准控制模型输出的研究者使用。
大语言模型的行为精准控制对复杂应用至关重要。然而,现有方法常伴随高训练成本、缺乏自然语言可控性或牺牲语义连贯性。为此,我们提出WASD(unWeaving Actionable Sufficient Directives)框架,通过识别生成特定标记的充分神经条件来解释模型行为。该方法将候选条件表示为神经元激活谓词,并迭代搜索在输入扰动下仍能保证当前输出的最小集合。在Gemma-2-2B模型上针对SST-2和CounterFact的实验表明,该方法生成的解释比传统归因图更稳定、准确且简洁。此外,跨语言生成控制案例研究验证了WASD在行为控制中的实际有效性。
原文摘要 · Abstract (English)
Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often incur high training costs, lack natural language controllability, or compromise semantic coherence. To bridge this gap, we propose WASD (unWeaving Actionable Sufficient Directives), a novel framework that explains model behavior by identifying sufficient neural conditions for token generation. Our method represents candidate conditions as neuron-activation predicates and iteratively searches for a minimal set that guarantees the current output under input perturbations. Experiments on SST-2 and CounterFact with the Gemma-2-2B model demonstrate that our approach produces explanations that are more stable, accurate, and concise than conventional attribution graphs. Moreover, through a case study on controlling cross-lingual output generation, we validated the practical effectiveness of WASD in controlling model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。