4B和9B参数模型专攻长文本,数学编码表现强。
xGen-small Technical Report
- 分阶段预训练+质量退火,支持128k上下文长度
- 在数学与编程任务中表现优异,长文本基准领先
- 适合需要处理长文档的开发者与研究者
我们提出xGen-small,一个由40亿和90亿参数组成的Transformer解码器系列,专为长上下文应用优化。其垂直整合的流程包括领域均衡、频率感知的数据筛选;多阶段预训练结合质量退火与上下文长度扩展至128k token;以及通过监督微调、偏好学习和在线强化学习进行针对性后训练。xGen-small在多种任务中表现出色,尤其在数学和代码生成领域,同时在长上下文基准测试中表现卓越。
原文摘要 · Abstract (English)
We introduce xGen-small, a family of 4B and 9B Transformer decoder models optimized for long-context applications. Our vertically integrated pipeline unites domain-balanced, frequency-aware data curation; multi-stage pre-training with quality annealing and length extension to 128k tokens; and targeted post-training via supervised fine-tuning, preference learning, and online reinforcement learning. xGen-small delivers strong performance across various tasks, especially in math and coding domains, while excelling at long context benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。