arXiv:2605.24414cs.AI2026-05

JT-Safe-V2通过融合世界知识提升模型安全与智能,适合企业级代理应用。

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

论文配图:JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data
图 1 · 摘自论文原文
  • 用世界上下文数据增强预训练,实现安全与智能联合优化。
  • 在通用与安全基准上达到顶尖水平,推理成本降低超30%。
  • 提出可追溯的Safe-MoMA框架,适合企业级智能体部署。

我们提出JT-Safe-V2,一种旨在提升基础模型安全性与可信度的大语言模型,延续并扩展了之前的JT-Safe模型,迈向更全面的安全设计范式。该模型通过多项创新实现通用智能与安全设计的联合优化:在预训练阶段引入上下文世界知识,采用高置信度预训练流程,并结合安全强化后训练机制,以支持面向企业应用的智能体能力。基于这些安全增强的基础模型,我们提出Safe-MoMA(安全的模型与智能体混合体)框架,通过协调部署多个模型与智能体,实现可追溯且高效的推理。大量评估表明,JT-Safe-V2在通用智能与安全基准测试中均达到当前最优性能。此外,相比使用最大独立模型基线,Safe-MoMA将推理成本降低超过30%,同时保持相当的性能水平。为促进未来对安全设计基础模型的研究,我们公开发布后训练后的JT-Safe-V2-35B模型检查点。

原文摘要 · Abstract (English)

We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe model toward a more comprehensive safety-by-design paradigm. JT-Safe-V2 emphasizes the joint optimization of general intelligence and safety-by-design through several key innovations: enriching pre-training data with contextual world knowledge, high-certainty pre-training procedures, and safety strengthening post-training mechanisms for enterprise-oriented agentic capabilities. Building on these safety-enhanced foundation models, we propose Safe-MoMA (Safe Mixture of Models and Agents), a framework that enables traceable and efficient inference through the orchestrated deployment of multiple models and agents. Extensive evaluations demonstrate that JT-Safe-V2 achieves state-of-the-art performance across both general intelligence and safety benchmarks. Moreover, Safe-MoMA reduces inference costs by more than 30\% compared to using the largest standalone model baseline while maintaining comparable performance. To facilitate future research on safety-by-design foundation models, we publicly release the post-trained JT-Safe-V2-35B model checkpoint.

安全模型智能体大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。