arXiv:2509.19762cs.AI2025-09

优化推理流程,让小模型表现超越大模型。

The Conductor and the Engine: A Path Towards Co-Designed Reasoning

  • 设计新推理工作流,协同模型能力与调度框架。
  • 小模型在任务上超越数倍大小的大型模型。
  • 适合追求高效推理的小模型研究者使用。

现代大模型推理依赖大量运行时计算,由模型内部训练和外部智能体编排共同驱动。然而,这种协同常因模型冗余和指令理解差而造成算力浪费。我们分析了能力与成本之间的权衡关系,提出一种优化的推理工作流(\ ext{cepo}),使小型开源模型的表现优于体积数倍于自身的模型。我们将开源该工作流以推动后续研究。本工作展示了将调度框架与底层模型能力协同设计的清晰路径,从而在中小规模模型中实现强大推理能力。

原文摘要 · Abstract (English)

Modern LLM reasoning relies on extensive test-time computation, driven by internal model training and external agentic orchestration. However, this synergy is often inefficient, as model verbosity and poor instruction following lead to wasted compute. We analyze this capability-cost trade-off and introduce an optimized reasoning workflow (\cepo) that empowers smaller open-source models to outperform models multiple times their size. We will open-source this workflow to enable further research. Our work demonstrates a clear path toward co-designing orchestration frameworks with the underlying model capabilities to unlock powerful reasoning in small-to-medium sized models.

推理优化小模型协同设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。