arXiv:2510.17918cs.CLcs.AI2025-10

通过注入世界上下文增强预训练数据,提升大模型安全可信度。

JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs

  • 用真实世界上下文重构预训练数据,使其更贴近现实
  • 在1.5万亿上下文增强数据上继续预训练,安全可信度提升1.79%
  • 适合关注模型安全、可信推理的研究者与工业应用开发者

大语言模型的幻觉和可信度问题是行业共同挑战。尽管后训练与推理阶段已有诸多改进,但根源在于预训练阶段——数据本身存在事实错误、逻辑不一致及分布偏差,且缺乏真实世界知识锚定。本文提出将预训练数据与真实世界上下文结合,构建‘世界上下文数据’(DWC),强调原始数据的时空背景与实际作用。我们基于早期的JT-35B-Base模型,在1.5万亿个DWC tokens上继续预训练,并设计后训练流程激活其潜力。相较于同规模的Qwen模型,JT-Safe-35B在安全与可信评估基准上平均提升1.79%,且仅使用6.2万亿预训练令牌。

原文摘要 · Abstract (English)

The hallucination and credibility concerns of large language models (LLMs) are global challenges that the industry is collectively addressing. Recently, a significant amount of advances have been made on post-training and inference techniques to mitigate these challenges. However, it is widely agreed that unsafe and hallucinations of LLMs intrinsically originate from pre-training, involving pre-training data and the next-token prediction learning mechanism. In this paper, we focus on enhancing pre-training data to improve the trustworthiness and safety of LLMs. Since the data is vast, it's almost impossible to entirely purge the data of factual errors, logical inconsistencies, or distributional biases. Moreover, the pre-training data lack grounding in real-world knowledge. Each piece of data is treated as a sequence of tokens rather than as a representation of a part of the world. To overcome these issues, we propose approaches to enhancing our pre-training data with its context in the world and increasing a substantial amount of data reflecting industrial scenarios. We argue that most source data are created by the authors for specific purposes in a certain spatial-temporal context. They have played a role in the real world. By incorporating related world context information, we aim to better anchor pre-training data within real-world scenarios, thereby reducing uncertainty in model training and enhancing the model's safety and trustworthiness. We refer to our Data with World Context as DWC. We continue pre-training an earlier checkpoint of JT-35B-Base with 1.5 trillion of DWC tokens. We introduce our post-training procedures to activate the potentials of DWC. Compared with the Qwen model of a similar scale, JT-Safe-35B achieves an average performance improvement of 1.79% on the Safety and Trustworthy evaluation benchmarks, while being pretrained with only 6.2 trillion tokens.

大模型安全预训练优化可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。