arXiv:2502.12446cs.CLcs.AI2025-02ACL被引 40

让大模型同时优化多个冲突属性,像调音一样精准控制输出质量。

Multi-Attribute Steering of Language Models via Targeted Intervention

  • 通过学习稀疏正交的引导向量,在生成时逐令牌干预模型行为。
  • 在问答任务中平均准确率提升3%,生成任务胜率超基线55.82%。
  • 适合需要多维度调控输出质量的场景,如客服、内容生成等。

推理时干预(ITI)已成为一种无需更新参数即可引导大语言模型行为的方法。然而现有方法难以处理多属性冲突场景,如提升帮助性的同时降低毒性。为此,我们提出多属性定向引导(MAT-Steer),一种针对多属性的细粒度令牌级干预框架。该方法利用对齐目标,使模型对不良输出的内部表征向理想输出靠拢,并强制不同属性间的引导向量保持稀疏与正交,从而缓解属性间冲突。我们在两类任务中评估:(i)问答任务中平衡真实性、偏见与毒性;(ii)生成任务中同步提升帮助性、正确性与连贯性。MAT-Steer在两类任务上均优于现有ITI与参数高效微调方法(如在问答任务中平均准确率提升3%,生成任务胜率达55.82%)。

原文摘要 · Abstract (English)

Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.g., improving helpfulness) by intervening on token representations without costly updates to the LLM's parameters. However, existing ITI approaches fail to scale to multi-attribute settings with conflicts, such as enhancing helpfulness while also reducing toxicity. To address this, we introduce Multi-Attribute Targeted Steering (MAT-Steer), a novel steering framework designed for selective token-level intervention across multiple attributes. MAT-Steer learns steering vectors using an alignment objective that shifts the model's internal representations of undesirable outputs closer to those of desirable ones while enforcing sparsity and orthogonality among vectors for different attributes, thereby reducing inter-attribute conflicts. We evaluate MAT-Steer in two distinct settings: (i) on question answering (QA) tasks where we balance attributes like truthfulness, bias, and toxicity; (ii) on generative tasks where we simultaneously improve attributes like helpfulness, correctness, and coherence. MAT-Steer outperforms existing ITI and parameter-efficient fine-tuning approaches across both task types (e.g., 3% average accuracy gain across QA tasks and 55.82% win rate against the best ITI baseline).

模型引导多属性优化推理干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。