通过抽象激活空间提升大模型形式推理的抗干扰能力
Abstract Activation Spaces for Content-Invariant Reasoning in Large Language Models
- 构建抽象与具体命题对,定义激活空间以分离语义与结构
- 用轻量级抽象器在多层干预中引导推理路径,减少内容偏差
- 跨语言测试验证其有效性,适合需要严谨逻辑的场景
大型语言模型在三段论推理中常因语义合理性误判形式有效性,这种内容效应持续存在,即使生成分步解释也未能缓解。为此,提出基于激活空间的抽象引导推理框架,通过配对含语义内容与抽象形式的三段论,利用模型在抽象输入上的激活定义抽象推理空间。设计轻量级抽象器,从含内容的残差流状态预测与该空间对齐的表示,并在前向传播中多层融合这些预测。以跨语言迁移为测试场景,结果表明该方法显著降低由语义驱动的错误,提升对形式有效性敏感的性能。研究证明激活层面的抽象是增强大模型形式推理鲁棒性、抵抗语义干扰的可扩展机制。
原文摘要 · Abstract (English)
Large Language Models (LLMs) often struggle with deductive judgment in syllogistic reasoning, systematically conflating semantic plausibility with formal validity a phenomenon known as content effect. This bias persists even when models generate step-wise explanations, indicating that intermediate rationales may inherit the same semantic shortcuts that affect answers. Recent approaches propose mitigating this issue by increasing inference-time structural constraints, either by encouraging abstract intermediate representations or by intervening directly in the model's internal computations; however, reliably suppressing semantic interference remains an open challenge. To make formal deduction less sensitive to semantic content, we introduce a framework for abstraction-guided reasoning that explicitly separates structural inference from lexical semantics. We construct paired content-laden and abstract syllogisms and use the model's activations on abstract inputs to define an abstract reasoning space. We then learn lightweight Abstractors that, from content-conditioned residual-stream states, predict representations aligned with this space and integrate these predictions via multi-layer interventions during the forward pass. Using cross-lingual transfer as a test bed, we show that abstraction-aligned steering reduces content-driven errors and improves validity-sensitive performance. Our results position activation-level abstraction as a scalable mechanism for enhancing the robustness of formal reasoning in LLMs against semantic interference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。