arXiv:2509.24147cs.CLcs.AI2025-09被引 3

用大模型分析大模型的思维模式,区分不同推理模型的思考习惯。

Your thoughts tell who you are: Characterize the reasoning patterns of LRMs

  • 用生成式模型自动提炼推理过程中的特征并分类
  • 80%-100%准确区分12个开源推理模型的思考路径
  • 发现小模型模仿大模型思维可提升3.3%-5.7%准确率

当前对大型推理模型(LRMs)的比较多聚焦于任务准确率或推理长度等宏观指标,而不同模型是否以不同方式推理仍是未解问题。为此,我们提出LLM提出的开放分类法(LOT),利用生成式语言模型对比两个LRM的推理轨迹,并以自然语言描述其独特特征。随后,基于这些特征在各模型输出中的经验分布,建模其对来源模型的预测能力。在推理轨迹数据集上迭代该流程,生成可读的人类语言分类体系,揭示模型思考差异。我们将LOT应用于12个开源LRM在数学、科学和编程任务上的推理行为,识别出系统性差异,在模型规模、基础架构或目标领域不同的情况下,分类准确率达80%-100%。除分类外,LOT提供的自然语言分类体系还提供了模型思考方式的定性解释。最后,在案例研究中,我们将推理差异与性能关联:在测试时将小型Qwen3模型的推理风格调整为与最大版本一致,使其在GPQA上的准确率提升3.3%-5.7%。

原文摘要 · Abstract (English)

Current comparisons of large reasoning models (LRMs) focus on macro-level statistics such as task accuracy or reasoning length. Whether different LRMs reason differently remains an open question. To address this gap, we introduce the LLM-proposed Open Taxonomy (LOT), a classification method that uses a generative language model to compare reasoning traces from two LRMs and articulate their distinctive features in words. LOT then models how these features predict the source LRM of a reasoning trace based on their empirical distributions across LRM outputs. Iterating this process over a dataset of reasoning traces yields a human-readable taxonomy that characterizes how models think. We apply LOT to compare the reasoning of 12 open-source LRMs on tasks in math, science, and coding. LOT identifies systematic differences in their thoughts, achieving 80-100% accuracy in distinguishing reasoning traces from LRMs that differ in scale, base model family, or objective domain. Beyond classification, LOT's natural-language taxonomy provides qualitative explanations of how LRMs think differently. Finally, in a case study, we link the reasoning differences to performance: aligning the reasoning style of smaller Qwen3 models with that of the largest Qwen3 during test time improves their accuracy on GPQA by 3.3-5.7%.

模型对比推理分析思维模式Qwen3

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。