基于Llama的中文多模态模型,支持函数调用与视觉理解。
The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities
- 在Llama 3.2基础上继续预训练,强化繁体中文表达能力。
- 在繁体中文任务中,函数调用与图像理解性能领先同规模模型。
- 支持移动端部署,开源模型与应用,适合中文AI开发者使用。
Llama-Breeze2(简称Breeze2)是一套参数量为3B和8B的先进多模态语言模型,专为增强繁体中文语言表征而设计。基于Llama 3.2模型家族,在大规模语料上继续进行预训练,以提升繁体中文的语言与文化表达能力。除语言建模外,模型显著增强了函数调用与视觉理解能力。据我们所知,在不使用推理诱导提示的情况下,Breeze2是当前同规模模型中繁体中文函数调用与图像理解性能最强的。其有效性在台湾通用知识、指令遵循、长上下文、函数调用及视觉理解等任务上得到验证。所有Breeze2模型均以Llama 3.2社区许可公开发布,并展示了在移动端运行的模型应用,该应用亦已开源。
原文摘要 · Abstract (English)
Llama-Breeze2 (hereinafter referred to as Breeze2) is a suite of advanced multi-modal language models, available in 3B and 8B parameter configurations, specifically designed to enhance Traditional Chinese language representation. Building upon the Llama 3.2 model family, we continue the pre-training of Breeze2 on an extensive corpus to enhance the linguistic and cultural heritage of Traditional Chinese. In addition to language modeling capabilities, we significantly augment the models with function calling and vision understanding capabilities. At the time of this publication, as far as we are aware, absent reasoning-inducing prompts, Breeze2 are the strongest performing models in Traditional Chinese function calling and image understanding in its size class. The effectiveness of Breeze2 is benchmarked across various tasks, including Taiwan general knowledge, instruction-following, long context, function calling, and vision understanding. We are publicly releasing all Breeze2 models under the Llama 3.2 Community License. We also showcase the capabilities of the model running on mobile platform with a mobile application which we also open source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。