arXiv:2601.21414cs.AIcs.CL2026-01被引 4

通过动态插值融合直觉与深思模型,实现思考效率与深度的平衡。

System 1&2 Synergy via Dynamic Model Interpolation

  • 用参数插值动态调节模型思考方式,不需额外训练。
  • 在5个数学推理任务上比思考型模型更准且更快。
  • 适合需要高效精准推理的AI应用开发人员。

训练一个能灵活切换直觉式(System 1)与深思式(System 2)认知模式的统一语言模型仍具挑战,因两种认知模式间存在干扰。现有研究多聚焦于提升系统2模型的效率,但仅控制输出长度,未能触及本质。本文提出从能力控制出发,调控模型‘如何思考’而非‘产出什么’。我们利用已有指令微调与思维链检查点,通过动态参数插值实现无训练调节。初步实验发现线性插值可生成凸且单调的帕累托前沿,源于表征连续性与结构连通性。基于此,提出DAMI框架,根据查询估算推理强度λ(q),以配置认知深度。训练阶段采用偏好学习联合优化准确率与效率;零样本部署时则基于模型间认知差异设计置信度方法。五项数学推理基准测试表明,DAMI在保持高效的同时,精度超越思考型模型,有效结合了系统1的效率与系统2的推理深度。

原文摘要 · Abstract (English)

Training a unified language model that adapts between intuitive System 1 and deliberative System 2 remains challenging due to interference between their cognitive modes. Recent studies have thus pursued making System 2 models more efficient. However, these approaches focused on output control, limiting what models produce. We argue that this paradigm is misaligned: output length is merely a symptom of the model's cognitive configuration, not the root cause. In this work, we shift the focus to capability control, which modulates \textit{how models think} rather than \textit{what they produce}. To realize this, we leverage existing Instruct and Thinking checkpoints through dynamic parameter interpolation, without additional training. Our pilot study establishes that linear interpolation yields a convex, monotonic Pareto frontier, underpinned by representation continuity and structural connectivity. Building on this, we propose \textbf{DAMI} (\textbf{D}yn\textbf{A}mic \textbf{M}odel \textbf{I}nterpolation), a framework that estimates a query-specific Reasoning Intensity $λ(q)$ to configure cognitive depth. For training-based estimation, we develop a preference learning method encoding accuracy and efficiency criteria. For zero-shot deployment, we introduce a confidence-based method leveraging inter-model cognitive discrepancy. Experiments on five mathematical reasoning benchmarks demonstrate that DAMI achieves higher accuracy than the Thinking model while remaining efficient, effectively combining the efficiency of System 1 with the reasoning depth of System 2.

模型融合推理增强动态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。