arXiv:2508.10599cs.AI2025-08被引 14

通过正交子空间分离,实现大模型多属性精准控制。

Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models

  • 为每个属性分配正交子空间,减少干扰。
  • 混合使用专属与共享子空间,提升控制精度。
  • 按语义动态干预关键词,适合精细调优场景。

激活操控为直接调节大语言模型行为提供了新途径,但现有方法难以同时调控多个属性,常引发冲突和权衡。为此,我们提出多子空间表示操控(MSRS)框架,通过子空间表示微调实现高效多属性操控。MSRS通过为各属性分配正交子空间,将影响隔离在模型表示空间中,降低属性间干扰。该框架采用混合子空间组合策略:结合属性特有子空间以实现独特操控方向,同时引入共享子空间用于共性操控方向,并设计动态加权函数以高效融合各成分。推理阶段,MSRS引入基于令牌的操控机制,动态识别并干预最具语义相关性的令牌,实现细粒度行为调节。实验表明,MSRS显著缓解属性冲突,在多种属性上超越现有方法,并在多样化下游任务中表现出良好泛化能力。

原文摘要 · Abstract (English)

Activation steering offers a promising approach to controlling the behavior of Large Language Models by directly manipulating their internal activations. However, most existing methods struggle to jointly steer multiple attributes, often resulting in interference and undesirable trade-offs. To address this challenge, we propose Multi-Subspace Representation Steering (MSRS), a novel framework for effective multi-attribute steering via subspace representation fine-tuning. MSRS reduces inter-attribute interference by allocating orthogonal subspaces to each attribute, isolating their influence within the model's representation space. MSRS also incorporates a hybrid subspace composition strategy: it combines attribute-specific subspaces for unique steering directions with a shared subspace for common steering directions. A dynamic weighting function learns to efficiently integrate these components for precise control. During inference, MSRS introduces a token-level steering mechanism that dynamically identifies and intervenes on the most semantically relevant tokens, enabling fine-grained behavioral modulation. Experimental results show that MSRS significantly reduces attribute conflicts, surpasses existing methods across a range of attributes, and generalizes effectively to diverse downstream tasks.

大模型操控属性对齐子空间学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。