arXiv:2412.13702cs.CLcs.AI2024-12被引 19

面向泰语的开源多模态大模型,支持文本、视觉与语音处理。

Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models

  • 基于Llama 3和Qwen2持续预训练,融合英泰混合数据提升泰语能力。
  • 推出1到700亿参数的泰语文本模型,含安全分类器与多模态理解能力。
  • 适合泰国本地化应用开发,如文档理解、语音转写与文化适配生成。

本文介绍Typhoon 2系列泰语开源大语言模型,涵盖文本、视觉与音频模态。Typhoon2-Text基于Llama 3和Qwen2,通过英泰混合数据进行持续预训练,并采用后训练技术强化泰语表现,同时保留原始模型能力。发布从1到700亿参数的文本模型,提供基础版与指令微调版。为保障生成安全,推出针对泰语文化优化的Typhoon2-Safety分类器。Typhoon2-Vision提升泰语文档理解能力,同时保持图像描述等通用视觉性能。Typhoon2-Audio引入端到端语音转语音架构,可处理音频、语音与文本输入,生成文本与语音输出。

原文摘要 · Abstract (English)

This paper introduces Typhoon 2, a series of text and multimodal large language models optimized for the Thai language. The series includes models for text, vision, and audio. Typhoon2-Text builds on state-of-the-art open models, such as Llama 3 and Qwen2, and we perform continual pre-training on a mixture of English and Thai data. We employ post-training techniques to enhance Thai language performance while preserving the base models' original capabilities. We release text models across a range of sizes, from 1 to 70 billion parameters, available in both base and instruction-tuned variants. To guardrail text generation, we release Typhoon2-Safety, a classifier enhanced for Thai cultures and language. Typhoon2-Vision improves Thai document understanding while retaining general visual capabilities, such as image captioning. Typhoon2-Audio introduces an end-to-end speech-to-speech model architecture capable of processing audio, speech, and text inputs and generating both text and speech outputs.

泰语模型多模态开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。