arXiv:2507.18013cs.CL2025-07被引 11

TeleChat新系列通过优化训练策略,实现推理与代码能力显著提升。

Technical Report of TeleChat2, TeleChat2.5 and T1

  • 通过强化学习和持续预训练,提升模型推理与编码能力
  • T1-115B在数学与编程任务上超越GPT-4o等闭源模型
  • 公开发布35B与115B参数版本,支持快速推理与复杂任务

我们介绍最新的TeleChat系列模型:TeleChat2、TeleChat2.5和T1,相较于前代TeleChat有显著性能提升。尽管模型架构变化极小,但通过改进预训练和后训练策略,实现显著进步。TeleChat2在10万亿高质量多样令牌上进行预训练,随后经由监督微调(SFT)与直接偏好优化(DPO)增强能力。TeleChat2.5与T1进一步引入领域特定数据的持续预训练,并结合强化学习(RL),显著提升代码生成与数学推理表现。T1专为复杂推理设计,支持长链式思维(CoT),在数学与编程任务中表现优异。而TeleChat2.5注重速度,实现快速推理。两者均为115B参数的稠密Transformer架构,相较原版TeleChat在推理与通用任务上均有大幅提升。值得注意的是,T1-115B在多项指标上超越OpenAI o1-mini与GPT-4o等闭源模型。我们公开发布TeleChat2、TeleChat2.5与T1,包括35B与115B参数的后训练版本,赋能开发者与研究人员构建多样化应用。

原文摘要 · Abstract (English)

We introduce the latest series of TeleChat models: \textbf{TeleChat2}, \textbf{TeleChat2.5}, and \textbf{T1}, offering a significant upgrade over their predecessor, TeleChat. Despite minimal changes to the model architecture, the new series achieves substantial performance gains through enhanced training strategies in both pre-training and post-training stages. The series begins with \textbf{TeleChat2}, which undergoes pretraining on 10 trillion high-quality and diverse tokens. This is followed by Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) to further enhance its capabilities. \textbf{TeleChat2.5} and \textbf{T1} expand the pipeline by incorporating a continual pretraining phase with domain-specific datasets, combined with reinforcement learning (RL) to improve performance in code generation and mathematical reasoning tasks. The \textbf{T1} variant is designed for complex reasoning, supporting long Chain-of-Thought (CoT) reasoning and demonstrating substantial improvements in mathematics and coding. In contrast, \textbf{TeleChat2.5} prioritizes speed, delivering rapid inference. Both flagship models of \textbf{T1} and \textbf{TeleChat2.5} are dense Transformer-based architectures with 115B parameters, showcasing significant advancements in reasoning and general task performance compared to the original TeleChat. Notably, \textbf{T1-115B} outperform proprietary models such as OpenAI's o1-mini and GPT-4o. We publicly release \textbf{TeleChat2}, \textbf{TeleChat2.5} and \textbf{T1}, including post-trained versions with 35B and 115B parameters, to empower developers and researchers with state-of-the-art language models tailored for diverse applications.

大模型推理能力代码生成开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。