arXiv:2604.10335cs.CLcs.LG2026-04

根据题目难易动态分配专家,提升数学推理准确率

Adaptive Multi-Expert Reasoning via Difficulty-Aware Routing and Uncertainty-Guided Aggregation

  • 按题目难度和不确定性选择不同专家组合
  • 在GSM8K上达75.28%准确率,优于多数7B模型
  • 适合需要高鲁棒性数学推理的场景

大语言模型在数学推理基准上表现优异,但面对不同难度问题时性能波动较大。本文提出自适应多专家推理(AMR)框架,通过动态调整策略应对问题复杂度。一个灵活的路由系统分析题干,预测难度与不确定性,并引导可重构采样机制控制生成广度。三个专用专家生成候选答案,经多轮修正与定稿阶段优化。神经验证器评估答案正确性,聚类聚合技术结合共识与质量确定最终答案。在GSM8K数据集上,AMR仅使用原始训练数据即达到75.28%准确率,超越多数基于合成数据训练的同类7B模型,证明基于难度的路由与不确定性驱动聚合能有效提升数学推理模型的鲁棒性。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate strong performance in math reasoning benchmarks, but their performance varies inconsistently across problems with varying levels of difficulty. This paper describes Adaptive Multi-Expert Reasoning (AMR), a framework that focuses on problem complexity by reasoning with dynamically adapted strategies. An agile routing system that focuses on problem text predicts problems' difficulty and uncertainty and guides a reconfigurable sampling mechanism to manage the breadth of generation. Three specialized experts create candidate responses, which are modified during multiple correction and finalization phases. A neural verifier assesses the correctness of responses, while a clustering-based aggregation technique identifies the final candidate answer based on a combination of consensus and answer quality. When evaluated on the GSM8K dataset, AMR achieved 75.28% accuracy while only using the original training data. This result outperformed the majority of comparable 7B models that were trained on synthetic data. This showcases that models using difficulty-based routing and uncertainty-driven aggregation are efficient and effective in improving math reasoning models' robustness.

数学推理多专家自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。