arXiv:2606.26108cs.CL2026-06

大模型更擅长约束引导推理,能更好识别并利用规则排除错误路径。

Where Larger Models Excel: The Primacy of Constraint-Guided Reasoning

论文配图:Where Larger Models Excel: The Primacy of Constraint-Guided Reasoning
图 1 · 摘自论文原文
  • 通过自动化框架分析大中小模型的推理差异
  • 大模型在数学物理等任务上平均领先6%-7%且稳定
  • 适合关注大模型推理机制与提升方向的研究者

大语言模型在推理基准测试中持续优于小模型,但其背后的具体推理差异仍不明确。在数学、物理、化学和编程多个基准上,我们观察到稳定的性能差距:平均而言,Qwen3-32B 比 Qwen3-8B 高出 6.43%,GPT-OSS-120B 比 GPT-OSS-20B 高出 7.38%。为探究这些优势背后的推理机制,我们开发了 AdvCluster——一个自动化框架,可识别大模型具有稳定优势的问题,从成对的推理轨迹中提取细粒度优势描述,并通过语义聚类与评审模型指导的量化评估进行筛选。分析揭示了一套系统性的大模型推理优势分类体系,涵盖跨领域共性优势及特定领域的专属优势。核心主题是‘约束引导推理’:大模型更善于识别显性和隐性约束,将其组织为结构化推理流程,并用于排除不可行路径和验证中间步骤。

原文摘要 · Abstract (English)

Larger language models consistently outperform smaller ones on reasoning benchmarks, yet the reasoning differences underlying this gap remain underexplored. Across benchmarks in mathematics, physics, chemistry, and programming, we observe stable performance gaps: averaged over datasets, Qwen3-32B outperforms Qwen3-8B by 6.43%, while GPT-OSS-120B exceeds GPT-OSS-20B by 7.38%. To study the reasoning differences behind these gains, we develop AdvCluster, an automated framework that identifies questions where the larger model shows a stable advantage, extracts fine-grained advantage descriptions from paired reasoning traces produced by larger and smaller models, and organizes them through semantic clustering with quantitative evaluation and selection guided by a reviewer model. Our analysis yields a systematic taxonomy of larger model reasoning advantages, spanning both common advantages that recur across domains and specialized advantages associated with particular domains. Across these patterns, a recurring theme is Constraint-Guided Reasoning: larger models are better at identifying explicit and implicit constraints, organizing them into structured reasoning, and using them to rule out infeasible paths and verify intermediate steps.

大模型推理约束推理模型对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。