arXiv:2503.10814cs.CL2025-03综述被引 33

系统梳理大模型推理方法,揭示其从语言生成到深度思考的跃迁路径。

Thinking Machines: A Survey of LLM based Reasoning Strategies

  • 按思维链、自洽性验证等策略分类归纳主流推理方法
  • 对比OpenAI O1、DeepSeek R1等模型在复杂任务上的推理表现
  • 适合关注AI可信性与通用智能的研究者和开发者

大型语言模型(LLMs)在语言任务中表现出色,其语言能力使其成为未来通用人工智能(AGI)竞争的核心。然而,近期研究(Valmeekam et al., 2024;Zecevic et al., 2023;Wu et al., 2024)指出,其语言能力与实际推理能力之间存在显著差距。通过引入推理机制,使LLMs与视觉语言模型(VLMs)具备思考与自我修正能力,是弥合该鸿沟的关键。推理能力对于解决复杂问题至关重要,也是建立对人工智能信任的必要步骤,从而推动其在医疗、金融、法律、国防、安全等敏感领域的部署。近年来,随着OpenAI O1和DeepSeek R1等强大推理模型的出现,推理能力赋予已成为LLM研究的核心议题。本文系统综述并比较现有推理技术,全面分析推理增强型语言模型的发展现状,并探讨当前挑战与关键发现。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are highly proficient in language-based tasks. Their language capabilities have positioned them at the forefront of the future AGI (Artificial General Intelligence) race. However, on closer inspection, Valmeekam et al. (2024); Zecevic et al. (2023); Wu et al. (2024) highlight a significant gap between their language proficiency and reasoning abilities. Reasoning in LLMs and Vision Language Models (VLMs) aims to bridge this gap by enabling these models to think and re-evaluate their actions and responses. Reasoning is an essential capability for complex problem-solving and a necessary step toward establishing trust in Artificial Intelligence (AI). This will make AI suitable for deployment in sensitive domains, such as healthcare, banking, law, defense, security etc. In recent times, with the advent of powerful reasoning models like OpenAI O1 and DeepSeek R1, reasoning endowment has become a critical research topic in LLMs. In this paper, we provide a detailed overview and comparison of existing reasoning techniques and present a systematic survey of reasoning-imbued language models. We also study current challenges and present our findings.

大模型推理AGI可信AI模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。