构建高质量阿拉伯语对话数据集,用于大模型后训练
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
- 采用多模型集成翻译+质量过滤,提升阿拉伯语数据可信度
- 在真实对话场景中实现多轮推理与工具调用的高质量数据覆盖
- 适合需要高精度阿拉伯语能力的AI模型研发者
尽管社区已关注阿拉伯语预训练数据的获取,但缺乏大规模、多轮次且包含推理与工具调用的阿拉伯语数据集。直接翻译适用于预训练阶段,但后训练对数据质量要求更高,需更严格的筛选策略。本文提出SmolKalam,基于Smoltalk2的翻译框架,采用多模型集成翻译管道,结合质量过滤机制,并通过消融实验验证传统解码器仅用模型在阿拉伯语翻译中的有效技术。
原文摘要 · Abstract (English)
Although the community has tackled the acquisition of high-quality Arabic pretraining data, we still lack large-scale, multi-turn Arabic datasets that include reasoning and tool calling. Naive translation can work at the pretraining scale, but post-training demands much higher quality, which requires a stricter approach to dataset curation. In this work, we introduce SmolKalam, a translation of Smoltalk2 that uses a multi-model ensemble translation pipeline, applies quality filtering, and examines effective translation techniques for traditional decoder-only models through ablations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。