arXiv:2606.12234cs.CL2026-06被引 1

研究大模型控制中效果与流畅性的权衡,发现高效控制常损害生成质量。

On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study

  • 系统对比多种控制方法在注入与移除概念时的表现。
  • 高效控制方法常显著降低输出流畅性,且对指令微调模型效果差。
  • 简单提示和全量微调适合概念注入,但移除效果不佳。

控制大语言模型(LLMs)的输出是其可靠部署的核心挑战,但其中涉及的权衡关系尚不清晰。现有条件化方法多仅关注目标概念的注入或移除效果,忽视生成质量。本文系统研究了多种条件化方法在注入与移除场景下的表现。结果表明,高效的操控方法往往以严重牺牲流畅性为代价。此外,我们发现一个此前被忽略的关键交互:激活引导方法在指令微调模型上的效果远低于基础模型。相比之下,简单提示和全量监督微调在概念注入上表现良好,但在概念移除方面效果较差。最后,低成本文本指标与昂贵的LLM作为裁判评分高度相关,能有效揭示条件化方法的行为特征。

原文摘要 · Abstract (English)

Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are often evaluated with a narrow focus on their effectiveness at injecting or removing a target concept, neglecting generation quality. We systematically investigate a range of conditioning methods in both injection and removal scenarios. We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency. Furthermore, we identify a critical yet previously overlooked interaction with the training paradigm: activation steering methods are far less effective on instruction-tuned models than on their base counterparts. Simple prompting and full-fledged supervised fine-tuning, on the other hand, are viable options for concept injection, but are not as good at concept removal. Finally, cheaply computed textual metrics highly correlate to costly LLM-as-judge scores, and provide insights on the behavior of conditioning methods.

大模型控制效果权衡生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。