开源1200亿参数模型经优化,支持威尔士语并提升推理与指令遵循能力。
Jupiter-N Technical Report
- 用不确定性筛选轨迹增强智能体能力,结合合成数据对齐英国文化
- 在威尔士语任务上表现提升18分,终端使用和指令遵循能力分别增9.1和4.4
- 可复现的主权微调模板,适合关注本地化AI的国家或机构参考
我们提出Jupiter-N,一个基于完全开源的1200亿参数大模型Nemotron 3 Super的混合推理模型。目标包括:(1) 通过不确定性筛选的轨迹增强智能体能力;(2) 基于文化规范的合成数据实现英国文化对齐;(3) 通过平行语料库与大模型翻译对话支持威尔士语。数据筛选策略采用Forget-Me-Not框架,融合在线合成回放与离线任务数据,防止灾难性遗忘,并保留原模型的混合推理能力。实验显示,相较于Nemotron,Jupiter-N在威尔士语任务上提升18分(ARC-Easy)、5.25分(MMLU-Lite),终端使用能力提升9.1分(Terminal Bench 2),指令遵循能力提升4.4分(IFBench),同时保持原有模型能力。本工作提供可复现的主权后训练范式:替换文化知识、机构语料与目标语言,即可适配任何国家。所有模型权重与后训练数据均以开放许可公开。
原文摘要 · Abstract (English)
We present Jupiter-N, a hybrid reasoning model post-trained from Nemotron 3 Super, a fully open-source 120 billion parameter LLM. We target three objectives: (1) agentic capability via uncertainty-curated trajectories; (2) UK cultural alignment via synthetic data grounded in cultural norms; and (3) Welsh language support via parallel corpora and LLM-translated Welsh conversations. Our data curation strategy carefully preserves the base model's capabilities: using our Forget-Me-Not framework, we mix on-policy synthetic replay with off-policy task data to mitigate catastrophic forgetting, and include a mixture of reasoning and non-reasoning traces to maintain Nemotron's hybrid reasoning ability. Jupiter-N achieves standout gains over Nemotron in Welsh (+18 on ARC-Easy, +5.25 on MMLU-Lite), terminal-use (+9.1 on Terminal Bench 2) and instruction following (+4.4 on IFBench), while retaining the base model capabilities. We frame this work as a reproducible template for sovereign post-training: substituting cultural knowledge, institutional corpora, and target languages produces an equivalent pipeline for any country. All model weights and all post-training datasets are publicly released under open licences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。