动态调整模型干预强度,让生成文本既可控又自然。
In-Distribution Steering: Balancing Control and Coherence in Language Model Generation
- 根据输入在表征空间中的分布距离,自适应调节干预强度。
- 在分类任务中保持高准确率,生成文本无崩溃且连贯。
- 适合需要稳定可控生成的真实场景应用。
激活控制方法通过在推理时修改大语言模型的内部激活来调控其行为。然而,现有方法多采用固定强度的干预,导致控制不足或干扰过度,降低文本合理性和连贯性。本文提出分布内控制(In-Distribution Steering, IDS),根据输入在表征空间中的分布位置动态调整干预强度,实现自适应干预与生成稳定性。实验表明,IDS在分类任务中取得优异准确率,同时生成文本保持连贯且无坍塌现象,特别适用于真实应用场景。
原文摘要 · Abstract (English)
Activation steering methods control large language model (LLM) behavior by modifying internal activations at inference time. However, most existing activation steering methods rely on a fixed steering strength, leading to either insufficient control or unadapted intervention that degrades text plausibility and coherence. We introduce In-Distribution Steering (IDS), a novel method that adapts steering strength based on the input data distribution in representation space. IDS dynamically adjusts interventions according to how far a given input lies within the distribution, enabling adaptive intervention and generation stability during text generation. Experiments demonstrate that IDS achieves strong accuracy on classification tasks while producing coherent text without collapse, making IDS particularly well suited for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。