arXiv:2512.02556cs.CL2025-12被引 732

DeepSeek-V3.2用高效注意力和强化学习,让开源大模型推理与工具使用能力逼近顶尖水平。

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

  • 采用稀疏注意力机制,降低计算开销同时保持长文本处理能力。
  • 高算力版本超越GPT-5,在数学与信息学奥赛中获金牌。
  • 构建大规模智能体任务生成流水线,提升复杂环境下的指令遵循能力。

我们介绍 DeepSeek-V3.2,一款在计算效率、推理能力和智能体表现上均达到领先水平的开源大语言模型。关键技术突破包括:(1) 深度稀疏注意力(DSA):一种高效的注意力机制,在长上下文场景下显著降低计算复杂度,同时保持模型性能;(2) 可扩展的强化学习框架:通过稳健的强化学习协议与大规模后训练算力,DeepSeek-V3.2 在性能上媲美 GPT-5。其高算力变体 DeepSeek-V3.2-Speciale 超越 GPT-5,推理能力与 Gemini-3.0-Pro 相当,在 2025 年国际数学奥林匹克(IMO)和国际信息学奥林匹克(IOI)中均取得金牌成绩;(3) 大规模智能体任务合成流水线:为将推理融入工具使用场景,开发了系统化的大规模训练数据生成方法,支持可扩展的智能体后训练,显著提升模型在复杂交互环境中的泛化性与指令遵循鲁棒性。

原文摘要 · Abstract (English)

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance in long-context scenarios. (2) Scalable Reinforcement Learning Framework: By implementing a robust reinforcement learning protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par with Gemini-3.0-Pro, achieving gold-medal performance in both the 2025 International Mathematical Olympiad (IMO) and the International Olympiad in Informatics (IOI). (3) Large-Scale Agentic Task Synthesis Pipeline: To integrate reasoning into tool-use scenarios, we developed a novel synthesis pipeline that systematically generates training data at scale. This methodology facilitates scalable agentic post-training, yielding substantial improvements in generalization and instruction-following robustness within complex, interactive environments.

大模型推理能力智能体开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。