arXiv:2604.02972cs.CL2026-04ACL被引 18

通过神经元混合机制提升大模型推理的可解释性与可控性

NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-Neurons

  • 识别关键神经元模式,构建统一的推理纠错框架
  • 在六项基准上最高提升27.0%性能,减少63.3%的令牌消耗
  • 适合需要可靠、可控推理的复杂任务应用

大型推理模型在复杂推理任务中取得显著进展,但存在三类固有缺陷:步骤内计算错误、步骤间震荡停滞、实例级过度思考。现有方法仅针对单一层面,且依赖强化学习导致可解释性差。本文通过白盒分析,发现与失败模式相关的关键神经元(神经元混合,MoN)及其波动规律,提出NeuReasoner框架,利用轻量MLP检测故障,并通过特殊标记触发自纠正机制,基于SFT训练实现可控修复。推理时插入特殊标记激活纠正行为。在六项基准、六种骨干模型(8B~70B)上对比九种基线,性能最高提升27.0%,令牌消耗降低19.6%~63.3%。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) have recently achieved remarkable success in complex reasoning tasks. However, closer scrutiny reveals persistent failure modes compromising performance and cost: I) Intra-step level, marked by calculation or derivation errors; II) Inter-step level, involving oscillation and stagnation; and III) Instance level, causing maladaptive over-thinking. Existing endeavors target isolated levels without unification, while their black-box nature and reliance on RL hinder explainability and controllability. To bridge these gaps, we conduct an in-depth white-box analysis, identifying key neurons (Mixture of Neurons, MoN) and their fluctuation patterns associated with distinct failures. Building upon these insights, we propose NeuReasoner, an explainable, controllable, and unified reasoning framework driven by MoN. Technically, NeuReasoner integrates lightweight MLPs for failure detection with a special token-triggered self-correction mechanism learned via SFT. During inference, special tokens are inserted upon failure detection to actuate controllable remedial behaviors. Extensive evaluations across six benchmarks, six backbone models (8B~70B) against nine competitive baselines, demonstrate that NeuReasoner achieves performance gains of up to 27.0% while reducing token consumption by 19.6% ~ 63.3%.

推理增强可解释性神经元分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。