arXiv:2505.03189cs.AIcs.HC2025-05被引 2

通过内部表示修改实现大模型行为控制,无需训练即可调参。

Patterns and Mechanisms of Contrastive Activation Engineering

  • 在推理时通过对比激活工程调整模型内部表征。
  • 80个样本后增益减弱,且可能被对抗输入逆转效果。
  • 适合需要快速定制大模型行为的开发者使用。

大型语言模型(LLM)的行为控制因其内在复杂性和不透明性而面临重大挑战。尽管微调等技术可改变模型行为,但通常需大量计算资源。近期提出的对比激活工程(CAE)方法可在推理阶段零成本地通过针对性修改内部表示来引导输出,具有引入灵活、任务专用行为调优新范式的潜力。本文分析了CAE在分布内与分布外设置下的表现,评估其缺陷,并初步建立有效部署指南。发现:1. CAE仅在分布内场景下可靠有效;2. 生成引导向量所用样本数超过80后收益递减;3. 引导向量易受对抗输入影响导致行为反转;4. 引导向量会降低模型整体困惑度;5. 模型越大,越能抵抗引导带来的性能退化。

原文摘要 · Abstract (English)

Controlling the behavior of Large Language Models (LLMs) remains a significant challenge due to their inherent complexity and opacity. While techniques like fine-tuning can modify model behavior, they typically require extensive computational resources. Recent work has introduced a class of contrastive activation engineering (CAE) techniques as promising approaches for steering LLM outputs through targeted modifications to their internal representations. Applied at inference-time with zero cost, CAE has the potential to introduce a new paradigm of flexible, task-specific LLM behavior tuning. We analyze the performance of CAE in in-distribution, out-of-distribution settings, evaluate drawbacks, and begin to develop comprehensive guidelines for its effective deployment. We find that 1. CAE is only reliably effective when applied to in-distribution contexts. 2. Increasing the number of samples used to generate steering vectors has diminishing returns at around 80 samples. 3. Steering vectors are susceptible to adversarial inputs that reverses the behavior that is steered for. 4. Steering vectors harm the overall model perplexity. 5. Larger models are more resistant to steering-induced degradation.

大模型控制激活工程推理调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。