arXiv:2510.10205cs.AI2025-10被引 15

提出像素级精准调控大模型生成行为的新方法,无需调参即可可靠对齐属性。

PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration

  • 基于双视角子空间学习,定位可干预的敏感位置
  • 通过闭式解自适应确定干预强度,避免全局调参
  • 适合需要可控生成且追求高可信度的应用场景

大语言模型在网页部署中的行为可靠性至关重要。激活值调控提供了一种无需微调即可对齐可信生成属性(如真实性)的方法。现有方法依赖粗糙启发式策略,缺乏对干预位置与强度的理论依据。为此,我们提出像素级激活注入框架 PIXEL,其核心是通过尾部平均与末标记双视角学习属性对齐子空间,并采用带约束的几何目标函数,以闭式解自适应选择干预强度,实现无需全局超参数调优的逐标记敏感性适配。PIXEL 进一步引入样本级正交残差校准,优化全局属性方向,并设计轻量级位置扫描流程识别最佳注入点。我们还提供了最小干预原则的表示层面保证,确保对齐可靠性。在多种模型和评估范式下,PIXEL 均显著提升属性对齐效果,同时保持模型通用能力,为大模型可控生成提供一种实用且原理严谨的方法。代码已开源:https://github.com/V1centNevwake/PIXEL-Adaptive-Steering

原文摘要 · Abstract (English)

Reliable behavior control is central to deploying large language models (LLMs) on the web. Activation steering offers a tuning-free route to align attributes (e.g., truthfulness) that ensure trustworthy generation. Prevailing approaches rely on coarse heuristics and lack a principled account of where to steer and how strongly to intervene. To this end, we propose Position-wise Injection with eXact Estimated Levels (PIXEL), a position-wise activation steering framework that, in contrast to prior work, learns a property-aligned subspace from dual views (tail-averaged and end-token) and selects intervention strength via a constrained geometric objective with a closed-form solution, thereby adapting to token-level sensitivity without global hyperparameter tuning. PIXEL further performs sample-level orthogonal residual calibration to refine the global attribute direction and employs a lightweight position-scanning routine to identify receptive injection sites. We additionally provide representation-level guarantees for the minimal-intervention rule, supporting reliable alignment. Across diverse models and evaluation paradigms, PIXEL consistently improves attribute alignment while preserving model general capabilities, offering a practical and principled method for LLMs' controllable generation. Our code is available at https://github.com/V1centNevwake/PIXEL-Adaptive-Steering

大模型控制激活调控可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。