让阿拉伯语大模型从预训练起就对齐,避免后期修正。
Alignment at Pre-training! Towards Native Alignment for Arabic LLMs
- 在预训练阶段直接引入对齐数据,实现原生对齐。
- 实验证明原生对齐显著提升模型性能与对齐稳定性。
- 开源了顶尖性能的阿拉伯语大模型,惠及社区。
大型语言模型(LLM)的对齐对于构建高效且安全的语言模型至关重要。传统方法通常在指令微调或强化学习阶段进行对齐,本文称之为‘后置对齐’。我们主张在预训练阶段即开展对齐,称为‘原生对齐’,旨在从源头防止内容偏离,而非依赖后期处理。该方法利用高度对齐的预训练数据,提升预训练模型的有效性与可用性。本研究聚焦阿拉伯语大模型的原生对齐应用,通过全面实验与消融研究,评估其对模型性能和对齐稳定性的提升效果。此外,我们发布了开源的阿拉伯语大模型,在多个基准测试中达到领先水平,为阿拉伯语大模型社区带来显著价值。
原文摘要 · Abstract (English)
The alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `post alignment'. We argue that alignment during the pre-training phase, which we term `native alignment', warrants investigation. Native alignment aims to prevent unaligned content from the beginning, rather than relying on post-hoc processing. This approach leverages extensively aligned pre-training data to enhance the effectiveness and usability of pre-trained models. Our study specifically explores the application of native alignment in the context of Arabic LLMs. We conduct comprehensive experiments and ablation studies to evaluate the impact of native alignment on model performance and alignment stability. Additionally, we release open-source Arabic LLMs that demonstrate state-of-the-art performance on various benchmarks, providing significant benefits to the Arabic LLM community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。