阿拉伯语生成式AI平台2.0,用极少数据实现顶尖性能。
Fanar 2.0: Arabic Generative AI Stack
- 以高质量数据为核心,持续预训练+模型融合提升性能
- 仅用1/8的训练量,阿拉伯语能力提升9.1分,英语能力增7.6分
- 支持阿拉伯语安全过滤、长音频识别、宗教内容生成等多场景
我们介绍Fanar 2.0,即卡塔尔以阿拉伯语为核心的生成式AI平台第二代。主权是首要设计原则:从数据管道到部署基础设施,所有组件均由哈马德·本·哈利法大学奎西研究创新中心(QCRI)自主设计与运营。该平台在资源受限条件下实现卓越成果:项目运行于256块NVIDIA H100 GPU,而阿拉伯语仅占网络数据约0.5%,尽管有4亿母语使用者。Fanar 2.0采用数据质量优先、定向持续预训练与模型融合策略,在约束下取得显著进步。核心为Fanar-27B,基于Gemma-3-27B基线,通过三个数据配方的1200亿高质量令牌持续预训练。虽使用预训练令牌数仅为Fanar 1.0的1/8,却在多项基准上实现显著提升:阿拉伯语知识+9.1分,语言能力+7.3分,方言能力+3.5分,英文能力+7.6分。除核心大模型外,平台还新增多项功能:FanarGuard为40亿参数双语内容安全过滤器;语音系列Aura引入支持小时级音频的长时序自动语音识别模型;Oryx视觉系列增强阿拉伯语感知图像与视频理解及文化语境生成能力;代理工具调用框架支持多步工作流;Fanar-Sadiq采用多智能体架构生成伊斯兰相关内容;Fanar-Diwan实现古典阿拉伯诗歌生成;FanarShaheen提供双语翻译能力;重新设计的多层编排器通过意图感知路由与纵深安全验证协调各组件。总体而言,Fanar 2.0证明主权且资源受限的AI开发可产出媲美大规模系统的成果。
原文摘要 · Abstract (English)
We present Fanar 2.0, the second generation of Qatar's Arabic-centric Generative AI platform. Sovereignty is a first-class design principle: every component, from data pipelines to deployment infrastructure, was designed and operated entirely at QCRI, Hamad Bin Khalifa University. Fanar 2.0 is a story of resource-constrained excellence: the effort ran on 256 NVIDIA H100 GPUs, with Arabic having only ~0.5% of web data despite 400 million native speakers. Fanar 2.0 adopts a disciplined strategy of data quality over quantity, targeted continual pre-training, and model merging to achieve substantial gains within these constraints. At the core is Fanar-27B, continually pre-trained from a Gemma-3-27B backbone on a curated corpus of 120 billion high-quality tokens across three data recipes. Despite using 8x fewer pre-training tokens than Fanar 1.0, it delivers substantial benchmark improvements: Arabic knowledge (+9.1 pts), language (+7.3 pts), dialects (+3.5 pts), and English capability (+7.6 pts). Beyond the core LLM, Fanar 2.0 introduces a rich stack of new capabilities. FanarGuard is a state-of-the-art 4B bilingual moderation filter for Arabic safety and cultural alignment. The speech family Aura gains a long-form ASR model for hours-long audio. Oryx vision family adds Arabic-aware image and video understanding alongside culturally grounded image generation. An agentic tool-calling framework enables multi-step workflows. Fanar-Sadiq utilizes a multi-agent architecture for Islamic content. Fanar-Diwan provides classical Arabic poetry generation. FanarShaheen delivers LLM-powered bilingual translation. A redesigned multi-layer orchestrator coordinates all components through intent-aware routing and defense-in-depth safety validation. Taken together, Fanar 2.0 demonstrates that sovereign, resource-constrained AI development can produce systems competitive with those built at far greater scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。