arXiv:2505.12189cs.AIcs.CL2025-05AAAI被引 39

通过激活调控减轻大模型推理中的内容偏见,提升逻辑判断准确性。

Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering

  • 在推理时调节模型内部激活,区分形式逻辑与内容合理性
  • 动态条件调控使未响应模型的逻辑准确率提升15%绝对值
  • 方法对提示变化鲁棒,且不影响多语言能力,适合实际部署

大型语言模型在推理中存在内容偏见,常将内容合理性误判为形式逻辑有效性,导致关键领域错误推断。本文研究通过激活调控这一推理时技术缓解此类偏见。首先定位负责形式推理与内容合理性的层,再在可控的三段论任务中测试激活调控效果。实验表明,对比性调控方法能线性控制内容偏见,但静态方法无法消除所有模型的偏差。为此提出基于kNN的细粒度条件调控方法(K-CAST),实现动态参数调整,显著降低未响应模型的偏见,使形式推理准确率最高提升15%绝对值。此外,该方法对提示变化具有鲁棒性,对多语言建模能力影响极小,并可部分推广至其他推理任务。实证显示,激活级干预是一种可扩展的推理时策略,有助于提升大模型系统性与无偏推理能力。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit reasoning biases, often conflating content plausibility with formal logical validity. This can lead to wrong inferences in critical domains, where plausible arguments are incorrectly deemed logically valid or vice versa. This paper investigates how content biases on reasoning can be mitigated through activation steering, an inference-time technique that modulates internal activations. Specifically, after localising the layers responsible for formal and plausible inference, we investigate activation steering on a controlled syllogistic reasoning task, designed to disentangle formal validity from content plausibility. An extensive empirical analysis reveals that contrastive steering methods consistently support linear control over content biases. However, a static approach is insufficient to debias all the tested models. We then investigate how to control content effects by dynamically determining the steering parameters through fine-grained conditional methods. By introducing a novel kNN-based conditional approach (K-CAST), we demonstrate that conditional steering can effectively reduce biases on unresponsive models, achieving up to 15% absolute improvement in formal reasoning accuracy. Finally, we found that steering for content effects is robust to prompt variations, incurs minimal side effects on multilingual language modeling capabilities, and can partially generalize to different reasoning tasks. In practice, we demonstrate that activation-level interventions offer a scalable inference-time strategy for enhancing the robustness of LLMs, contributing towards more systematic and unbiased reasoning capabilities.

推理偏见激活调控逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。