arXiv:2502.04043cs.LGcs.AI2025-02NAACL被引 5

无需探测器的低秩激活干预,提升大模型生成内容的真实性与质量

Probe-Free Low-Rank Activation Intervention

  • 用样本自适应的低秩非线性映射直接修改注意力层激活值
  • 在多个基础模型上显著提升生成与选择任务中的真实性和质量
  • 无需训练探测分类器,可高效求解优化问题,适合部署于实际系统

语言模型虽能生成看似准确连贯的内容,但常包含虚假或有害信息。推理阶段通过编辑隐藏激活值来引导模型生成更理想内容的方法展现出良好效果。现有方法多依赖激活探测器识别不良输出,再触发修正。本文提出无需探测器的FLORAIN方法,针对特定激活层的所有注意力头进行干预。该方法采用样本相关的非线性低秩映射作为干预函数,通过最小化修改后激活值与其在理想内容流形上的投影之间的距离进行训练。在特定流形与距离构造下,干预策略可通过求解光滑优化问题高效实现。实验结果表明,该方法在多个基础模型上均优于多个基线,在生成与多项选择任务中持续提升模型真实性与生成质量。

原文摘要 · Abstract (English)

Language models (LMs) can produce texts that appear accurate and coherent but contain untruthful or toxic content. Inference-time interventions that edit the hidden activations have shown promising results in steering the LMs towards desirable generations. Existing activation intervention methods often comprise an activation probe to detect undesirable generation, triggering the activation modification to steer subsequent generation. This paper proposes a probe-free intervention method FLORAIN for all attention heads in a specific activation layer. It eliminates the need to train classifiers for probing purposes. The intervention function is parametrized by a sample-wise nonlinear low-rank mapping, which is trained by minimizing the distance between the modified activations and their projection onto the manifold of desirable content. Under specific constructions of the manifold and projection distance, we show that the intervention strategy can be computed efficiently by solving a smooth optimization problem. The empirical results, benchmarked on multiple base models, demonstrate that FLORAIN consistently outperforms several baseline methods in enhancing model truthfulness and quality across generation and multiple-choice tasks.

模型干预真实性提升低秩映射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。