arXiv:2601.06403cs.CL2026-01被引 4

通过对比解码实现模型行为的连续调控,无需重新训练。

Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding

  • 用目标与默认系统提示的对数差异,分离出特定人格的行为信号。
  • 在五项基准测试中,严格准确率最高提升8.5%,拒绝率提高45个百分点。
  • 适合需要灵活控制模型角色、风格或能力的开发者使用。

大型语言模型在复杂指令上表现优异,但难以脱离其助手机能设定,因后训练阶段强化了固定先验,抗拒冲突指令。本文提出系统提示强度(system prompt strength)这一无需训练的方法,将提示遵循程度视为连续可控变量。通过对比目标提示与默认提示的对数输出,利用标量因子α放大仅属于目标人格的行为信号。在涵盖约束满足、行为控制、多元对齐、能力调节和风格控制的五个不同基准上,该方法均取得显著改进:IFEval上严格准确率最高提升8.5%,OffTopicEval拒绝率提升45个百分点,Prompt-Steering可调性提升13%。该方法使实践者可在不重训练的前提下动态调节模型行为。

原文摘要 · Abstract (English)

Large language models excel at complex instructions yet struggle to deviate from their helpful assistant persona, as post-training instills strong priors that resist conflicting instructions. We introduce system prompt strength, a training-free method that treats prompt adherence as a continuous control. By contrasting logits from target and default system prompts, we isolate and amplify the behavioral signal unique to the target persona by a scalar factor alpha. Across five diverse benchmarks spanning constraint satisfaction, behavioral control, pluralistic alignment, capability modulation, and stylistic control, our method yields substantial improvements: up to +8.5 strict accuracy on IFEval, +45pp refusal rate on OffTopicEval, and +13% steerability on Prompt-Steering. Our approach enables practitioners to modulate system prompt strength, providing dynamic control over model behavior without retraining.

提示控制行为调节零样本调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。