arXiv:2503.06072cs.CLcs.AI2025-03综述被引 55

系统梳理大模型后训练技术,揭示推理与对齐的关键进展

A Survey on Post-training of Large Language Models

  • 归纳五类后训练范式:微调、对齐、推理、效率优化与多模态融合
  • 对比ChatGPT到DeepSeek-R1的演进,展现推理能力跃升
  • 适合关注大模型优化、伦理对齐与应用落地的研究者阅读

大语言模型(LLMs)的兴起彻底改变了自然语言处理,广泛应用于对话系统到科学探索等领域。然而其预训练架构在特定场景中仍存在推理能力受限、伦理不确定性及领域性能不足等问题。为此,需通过后训练语言模型(PoLMs)弥补缺陷,如OpenAI-o1/o3和DeepSeek-R1(统称为大推理模型,LRMs)。本文首次全面综述PoLMs,系统梳理其五大核心范式:微调(提升任务精度)、对齐(确保伦理一致与人类偏好)、推理(克服奖励设计难题以实现多步推断)、效率(优化资源利用应对复杂性增长)、集成与适配(跨模态扩展并解决一致性问题)。从ChatGPT的对齐策略到DeepSeek-R1的推理创新,展示如何利用数据集缓解偏见、增强推理能力与领域适应性。贡献包括开创性地整合PoLM发展脉络、构建技术与数据集的结构化分类体系,并提出以LRMs为核心的未来研究战略。作为首份覆盖该领域的综述,本文凝聚近期进展,建立严谨学术框架,推动高精度、伦理鲁棒、跨域通用的大模型发展。

原文摘要 · Abstract (English)

The emergence of Large Language Models (LLMs) has fundamentally transformed natural language processing, making them indispensable across domains ranging from conversational systems to scientific exploration. However, their pre-trained architectures often reveal limitations in specialized contexts, including restricted reasoning capacities, ethical uncertainties, and suboptimal domain-specific performance. These challenges necessitate advanced post-training language models (PoLMs) to address these shortcomings, such as OpenAI-o1/o3 and DeepSeek-R1 (collectively known as Large Reasoning Models, or LRMs). This paper presents the first comprehensive survey of PoLMs, systematically tracing their evolution across five core paradigms: Fine-tuning, which enhances task-specific accuracy; Alignment, which ensures ethical coherence and alignment with human preferences; Reasoning, which advances multi-step inference despite challenges in reward design; Efficiency, which optimizes resource utilization amidst increasing complexity; Integration and Adaptation, which extend capabilities across diverse modalities while addressing coherence issues. Charting progress from ChatGPT's alignment strategies to DeepSeek-R1's innovative reasoning advancements, we illustrate how PoLMs leverage datasets to mitigate biases, deepen reasoning capabilities, and enhance domain adaptability. Our contributions include a pioneering synthesis of PoLM evolution, a structured taxonomy categorizing techniques and datasets, and a strategic agenda emphasizing the role of LRMs in improving reasoning proficiency and domain flexibility. As the first survey of its scope, this work consolidates recent PoLM advancements and establishes a rigorous intellectual framework for future research, fostering the development of LLMs that excel in precision, ethical robustness, and versatility across scientific and societal applications.

大模型后训练推理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。