不训练模型,用nPMI识别专家提升推理效率
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
- 用nPMI识别专注思考的'认知专家',动态调控推理过程
- 在DeepSeek-R1和Qwen3-235B上提升推理准确率与效率
- 轻量级设计,适合想优化推理但无训练资源的研究者
基于大规模推理模型(LRM)的混合专家(MoE)架构通过选择性激活专家,实现了结构化认知过程。然而,现有模型常存在过度思考和思考不足等认知效率问题。为此,我们提出一种无需额外训练的推理阶段引导方法RICE,利用归一化点互信息(nPMI)系统识别专门处理元层推理操作(如'<think>'标记)的'认知专家'。在主流MoE型推理模型(DeepSeek-R1与Qwen3-235B)上,针对严谨的定量与科学推理基准测试,实验表明该方法显著且一致地提升了推理准确率、认知效率及跨领域泛化能力。关键的是,其轻量级方案明显优于常见的提示工程与解码约束策略,同时保持模型通用指令遵循能力。结果表明,强化认知专家是一种有前景、实用且可解释的提升先进推理模型认知效率的方向。
原文摘要 · Abstract (English)
Mixture-of-Experts (MoE) architectures within Large Reasoning Models (LRMs) have achieved impressive reasoning capabilities by selectively activating experts to facilitate structured cognitive processes. Despite notable advances, existing reasoning models often suffer from cognitive inefficiencies like overthinking and underthinking. To address these limitations, we introduce a novel inference-time steering methodology called Reinforcing Cognitive Experts (RICE), designed to improve reasoning performance without additional training or complex heuristics. Leveraging normalized Pointwise Mutual Information (nPMI), we systematically identify specialized experts, termed ''cognitive experts'' that orchestrate meta-level reasoning operations characterized by tokens like ''<think>''. Empirical evaluations with leading MoE-based LRMs (DeepSeek-R1 and Qwen3-235B) on rigorous quantitative and scientific reasoning benchmarks demonstrate noticeable and consistent improvements in reasoning accuracy, cognitive efficiency, and cross-domain generalization. Crucially, our lightweight approach substantially outperforms prevalent reasoning-steering techniques, such as prompt design and decoding constraints, while preserving the model's general instruction-following skills. These results highlight reinforcing cognitive experts as a promising, practical, and interpretable direction to enhance cognitive efficiency within advanced reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。