Apertus推出全开源大模型,兼顾合规与多语言支持。
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
- 使用可复现的开放数据训练,严格遵守robots.txt并过滤敏感内容。
- 在15万亿tokens、1800+语言上训练,非英语数据占比约40%。
- 提供完整代码与工具链,适合研究者和开发者透明复现与扩展。
我们提出Apertus,一套完全开源的大语言模型套件,旨在解决当前开源模型生态中的两大系统性问题:数据合规性和多语言覆盖不足。与许多仅发布权重而无可复现数据管道或忽视内容版权的模型不同,Apertus模型仅在公开可用的数据上预训练,并回溯性地尊重`robots.txt`排除规则,过滤非授权、有害及个人身份信息内容。为降低记忆风险,预训练中采用Goldfish目标函数,显著抑制原文复现,同时保持下游任务性能。模型覆盖15万亿token、超过1800种语言,其中约40%预训练数据为非英语内容。以8B和70B规模发布,其在多语言基准测试中接近甚至超越现有完全开源模型的顶尖水平。除模型权重外,所有科研成果(包括数据处理脚本、检查点、评估工具和训练代码)均以宽松许可证开放,支持透明审计与持续拓展。
原文摘要 · Abstract (English)
We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingual representation. Unlike many prior models that release weights without reproducible data pipelines or regard for content-owner rights, Apertus models are pretrained exclusively on openly available data, retroactively respecting `robots.txt` exclusions and filtering for non-permissive, toxic, and personally identifiable content. To mitigate risks of memorization, we adopt the Goldfish objective during pretraining, strongly suppressing verbatim recall of data while retaining downstream task performance. The Apertus models also expand multilingual coverage, training on 15T tokens from over 1800 languages, with ~40% of pretraining data allocated to non-English content. Released at 8B and 70B scales, Apertus approaches state-of-the-art results among fully open models on multilingual benchmarks, rivalling or surpassing open-weight counterparts. Beyond model weights, we release all scientific artifacts from our development cycle with a permissive license, including data preparation scripts, checkpoints, evaluation suites, and training code, enabling transparent audit and extension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。