arXiv:2502.03618cs.LG2025-02ICML

用向量操作让模型按逻辑条件生成内容,解释性更强。

The Logical Implication Steering Method for Conditional Interventions on Transformer Generation

  • 通过向激活值添加概念向量实现生成控制
  • 可让模型在识别概念时自动触发指定行为
  • 适合需要透明推理的AI系统设计

预训练变压器模型的机制可解释性研究已充分支持「线性表征假说」,即高层次概念以激活空间中的向量形式编码。研究表明,通过向对应激活值中添加概念向量,可引导模型生成行为朝向特定概念。本文提出逻辑蕴含模型操控(LIMS)方法,利用该特性在模型中构建逻辑蕴含关系,使模型在识别任一给定概念时,能透明、可解释地触发预定生成行为。该方法通过将神经符号逻辑整合进预训练变压器模型,实现了人工精心设计的推理能力,拓展了模型的可控生成能力。

原文摘要 · Abstract (English)

The field of mechanistic interpretability in pre-trained transformer models has demonstrated substantial evidence supporting the ''linear representation hypothesis'', which is the idea that high level concepts are encoded as vectors in the space of activations of a model. Studies also show that model generation behavior can be steered toward a given concept by adding the concept's vector to the corresponding activations. We show how to leverage these properties to build a form of logical implication into models, enabling transparent and interpretable adjustments that induce a chosen generation behavior in response to the presence of any given concept. Our method, Logical Implication Model Steering (LIMS), unlocks new hand engineered reasoning capabilities by integrating neuro-symbolic logic into pre-trained transformer models.

逻辑推理可控生成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。