arXiv:2605.11906cs.CL2026-05

用神经元激活信号增强数学推理模型的偏好优化

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning

论文配图:YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning
图 1 · 摘自论文原文
  • 通过激活差值识别数学相关神经元,构建内部奖励信号
  • 在GSM8K上实验显示性能偶有提升,验证了内部信号有效性
  • 适合关注可解释性与细粒度训练的推理模型研究者

偏好优化已成为提升大语言模型推理能力的重要后训练方法。现有方法通常依赖外部构建的偏好数据,以优选和劣选响应作为样本级监督。然而,这些外部信号很少利用模型内部表征中蕴含的能力相关信息。在数学推理任务中,某些神经元组可能表现出与数学知识、符号操作或逻辑推理相关的激活模式。类似于反射性行为信号,这些内部激活可粗略指示模型是否调用了数学能力。本文提出YFPO(Yoked Feature Preference Optimization),一种针对数学推理的初步神经元引导偏好优化框架。YFPO首先使用AttnLRP识别数学相关神经元,再基于优选与劣选响应间其激活差值构建辅助奖励。该设计将内部神经元级信号融入外部偏好学习。我们在小型语言模型上以GSM8K为主要基准进行初步实验。结果表明,神经元级信号能与偏好优化交互,并在某些情况下提升推理性能,为更细粒度、可解释的推理导向后训练提供了有前景的方向。

原文摘要 · Abstract (English)

Preference optimization has become an important post-training paradigm for improving the reasoning abilities of large language models. Existing methods typically rely on externally constructed preference data, using preferred and dispreferred responses as sample-level supervision. However, such external signals rarely make explicit use of capability-related information contained in the model's internal representations. For mathematical reasoning, certain neuron groups may exhibit activation patterns associated with mathematical knowledge, symbolic manipulation, or logical reasoning. Similar to reflexive behavioral signals, these internal activations may provide a coarse indication of whether the model is engaging math-related capabilities.We introduce YFPO, short for Yoked Feature Preference Optimization, a preliminary neuron-guided preference optimization framework for mathematical reasoning. YFPO first uses AttnLRP to identify math-related neurons, and then constructs an auxiliary reward from their activation margin between preferred and dispreferred responses. This design augments external preference learning with internal neuron-level signals. We conduct preliminary experiments on a small-scale language model using GSM8K as the main benchmark. Results suggest that neuron-level signals can interact with preference optimization and occasionally improve reasoning performance, offering a promising direction for more fine-grained and interpretable reasoning-oriented post-training.

数学推理神经元分析偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。