arXiv:2602.08658cs.CL2026-02ACL被引 2

通过三种基本推理方式提升大模型跨领域泛化能力

Fundamental Reasoning Paradigms Induce Out-of-Domain Generalization in Language Models

  • 构建符号任务数据集,分别对应演绎、归纳、类比三种推理范式
  • 在真实自然语言任务上实现最高14.60的性能提升,显著增强泛化能力
  • 适合关注大模型逻辑推理与泛化能力的研究者和应用开发者

演绎、归纳和类比是人类逻辑思维的核心推理范式。尽管提升大语言模型(LLM)的推理能力已受到广泛关注,但这些基础范式如何促进模型泛化仍缺乏系统研究。本研究探讨了这三种核心范式对LLM推理行为的影响。我们首先收集了一个新的符号任务推理轨迹数据集,每个任务聚焦一种基本推理范式,以剥离具体世界知识。随后,我们实验了多种方法诱导这些推理技能,包括简单微调、增加模型深度,以及将密集模型转换为专家混合(MoE)结构。我们在完全由自然语言表述且包含真实世界知识的现实域外任务上全面评估了训练后的模型。结果表明,所提方法在真实任务上表现出强泛化性,性能提升最高达14.60。

原文摘要 · Abstract (English)

Deduction, induction, and abduction are fundamental reasoning paradigms, core for human logical thinking. Although improving Large Language Model (LLM) reasoning has attracted significant research efforts, the extent to which the fundamental paradigms induce generalization has yet to be systematically explored. In this study, we shed light on how the interplay between these core paradigms influences LLMs' reasoning behavior. To this end, we first collect a new dataset of reasoning trajectories from symbolic tasks, each targeting one of the three fundamental paradigms, to abstract from concrete world knowledge. Then, we investigate effective ways for inducing these skills into LLMs. We experiment with a battery of methods including simple fine-tuning, and more complex approaches to increase model depth, or transform a dense model to a mixture-of-experts. We comprehensively evaluate induced models on realistic out-of-domain tasks, that are entirely formulated in natural language and contain real-world knowledge. Our results reveal that our approach yields strong generalizability with substantial performance gains (up to $14.60$) across realistic tasks.

推理范式泛化能力大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。