38亿参数小模型实现强推理与多模态融合,性能媲美更大模型。
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- 用高质量合成数据训练,结合分组查询注意力提升效率。
- 在数学编程任务上超越同规模模型,语音识别排名第一。
- 适合资源受限场景的多模态应用,如移动设备部署。
我们推出Phi-4-Mini和Phi-4-Multimodal两款紧凑但强大的语言与多模态模型。Phi-4-Mini为38亿参数语言模型,基于高质量网络与合成数据训练,在数学与编码等复杂推理任务上表现优异,超越同规模开源模型,性能相当于两倍大小模型。其词汇量扩展至20万,支持多语言,并采用分组查询注意力以提升长序列生成效率。Phi-4-Multimodal将文本、视觉与语音/音频输入统一建模,通过LoRA适配器与模态专用路由实现无干扰多模态融合,支持(视觉+语言)、(视觉+语音)及(语音/音频)等多种输入组合。其语音模块仅4.6亿参数,却在OpenASR榜单中位列第一。该模型在多种任务上优于更大规模的视觉-语言与语音-语言模型。此外,对Phi-4-Mini进一步微调后,其推理能力达到或超越深求-7B与深求-8B等大模型水平。
原文摘要 · Abstract (English)
We introduce Phi-4-Mini and Phi-4-Multimodal, compact yet highly capable language and multimodal models. Phi-4-Mini is a 3.8-billion-parameter language model trained on high-quality web and synthetic data, significantly outperforming recent open-source models of similar size and matching the performance of models twice its size on math and coding tasks requiring complex reasoning. This achievement is driven by a carefully curated synthetic data recipe emphasizing high-quality math and coding datasets. Compared to its predecessor, Phi-3.5-Mini, Phi-4-Mini features an expanded vocabulary size of 200K tokens to better support multilingual applications, as well as group query attention for more efficient long-sequence generation. Phi-4-Multimodal is a multimodal model that integrates text, vision, and speech/audio input modalities into a single model. Its novel modality extension approach leverages LoRA adapters and modality-specific routers to allow multiple inference modes combining various modalities without interference. For example, it now ranks first in the OpenASR leaderboard to date, although the LoRA component of the speech/audio modality has just 460 million parameters. Phi-4-Multimodal supports scenarios involving (vision + language), (vision + speech), and (speech/audio) inputs, outperforming larger vision-language and speech-language models on a wide range of tasks. Additionally, we experiment to further train Phi-4-Mini to enhance its reasoning capabilities. Despite its compact 3.8-billion-parameter size, this experimental version achieves reasoning performance on par with or surpassing significantly larger models, including DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Llama-8B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。