arXiv:2602.04903cs.LG2026-02被引 1

特征操控虽能控制大模型行为,但会严重损害模型性能。

Mind the Performance Gap: Capability-Behavior Trade-offs in Feature Steering

  • 通过直接操纵内部表征实现行为控制,替代传统提示工程。
  • 在MMLU测试中,模型准确率从66%降至46%(Llama-8B), coherence 从4.62降至2.24。
  • 适合关注模型可控性与性能平衡的研究者或实际应用开发者。

特征操控作为一种通过直接修改内部表示来控制大语言模型行为的新兴方法,相比提示工程具有优势。然而其在真实场景中的有效性尚不明确,尤其存在与输出质量之间的潜在权衡。本文评估了Goodfire的Auto Steer方法,在171个MMLU问题上针对14个操控任务(涵盖无害与安全相关行为)进行测试,使用Llama-8B和Llama-70B模型,衡量准确性、连贯性与行为控制效果。结果表明,Auto Steer虽成功改变目标行为(在Llama-8B上得分为3.33,优于提示法的2.98;在Llama-70B上为3.57,优于3.10),但导致显著性能下降:准确率从66%降至46%(Llama-8B),87%降至73%(Llama-70B);连贯性从4.62降至2.24,4.94降至3.89。简单提示法在整体表现上更优。研究揭示当前特征操控方法在实际部署中面临能力-行为的根本性权衡,需在部署前实证评估。

原文摘要 · Abstract (English)

Feature steering has emerged as a promising approach for controlling LLM behavior through direct manipulation of internal representations, offering advantages over prompt engineering. However, its practical effectiveness in real-world applications remains poorly understood, particularly regarding potential trade-offs with output quality. We show that feature steering methods substantially degrade model performance even when successfully controlling target behaviors, a critical trade-off. Specifically, we evaluate Goodfire's Auto Steer against prompt engineering baselines across 14 steering queries (covering innocuous and safety-relevant behaviors) on 171 Massive Multitask Language Understanding (MMLU) questions using Llama-8B and Llama-70B, measuring accuracy, coherence, and behavioral control. Our findings show that Auto Steer successfully modifies target behaviors (achieving scores of 3.33 vs. 2.98 for prompting on Llama-8B and 3.57 vs. 3.10 on Llama-70B), but causes dramatic performance degradation: accuracy on the MMLU questions drops from 66% to 46% on Llama-8B and 87% to 73% on Llama-70B, with coherence falling from 4.62 to 2.24 and 4.94 to 3.89 respectively. Simple prompting achieves the best overall balance. These findings highlight limitations of current feature steering methods for practical deployment where task performance cannot be sacrificed. More broadly, our work demonstrates that mechanistic control methods face fundamental capability-behavior trade-offs that must be empirically characterized before deployment.

特征操控大模型控制性能权衡可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。