开源3.7B激活参数多模态大模型,专为企业任务优化
Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications
- 采用MoE架构,3.7B激活参数,40B总参数,专注企业场景
- 在RAG、表格理解等任务上表现领先,数学科学推理准确率接近顶尖模型
- 提出新强化学习算法,减少大模型过度思考,降低约一半至四分之三的生成成本
我们介绍Yuan3.0 Flash,一个开源的混合专家(MoE)多模态大语言模型,激活参数为3.7B,总参数达40B,专为提升企业任务性能而设计,同时保持通用任务上的竞争力。为解决大型推理模型中常见的过度思考现象,我们提出反射感知自适应策略优化(RAPO)算法,一种新型强化学习训练方法,可有效调控过度思考行为。在检索增强生成(RAG)、复杂表格理解、摘要生成等企业任务中,Yuan3.0 Flash表现持续领先。此外,其在数学、科学等领域的推理能力也表现出色,准确率与前沿模型相当,但仅需约1/4至1/2的平均生成token数。Yuan3.0 Flash已全面开源,以支持进一步研究和实际部署:https://github.com/Yuan-lab-LLM/Yuan3.0。
原文摘要 · Abstract (English)
We introduce Yuan3.0 Flash, an open-source Mixture-of-Experts (MoE) MultiModal Large Language Model featuring 3.7B activated parameters and 40B total parameters, specifically designed to enhance performance on enterprise-oriented tasks while maintaining competitive capabilities on general-purpose tasks. To address the overthinking phenomenon commonly observed in Large Reasoning Models (LRMs), we propose Reflection-aware Adaptive Policy Optimization (RAPO), a novel RL training algorithm that effectively regulates overthinking behaviors. In enterprise-oriented tasks such as retrieval-augmented generation (RAG), complex table understanding, and summarization, Yuan3.0 Flash consistently achieves superior performance. Moreover, it also demonstrates strong reasoning capabilities in domains such as mathematics, science, etc., attaining accuracy comparable to frontier model while requiring only approximately 1/4 to 1/2 of the average tokens. Yuan3.0 Flash has been fully open-sourced to facilitate further research and real-world deployment: https://github.com/Yuan-lab-LLM/Yuan3.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。