45亿参数小模型实现多模态通用处理,性能逼近顶尖水平
Towards Multi-Modal Mastery: A 4.5B Parameter Truly Multi-Modal Small Language Model
- 基于语言建模与多任务学习,构建跨文本图像视频音频的统一框架
- 在多个基准测试中表现接近顶尖水平,支持边缘设备部署
- 适合需要轻量化多模态能力的应用场景,如移动端AI
我们提出一种45亿参数的小型语言模型,可处理文本、图像、视频和音频等多种输入输出模态。尽管规模较小,该模型在多项任务中达到接近当前最优的性能,展现出多模态模型解决复杂现实问题的潜力。方法结合了语言建模与多任务学习的最新进展,构建了一个多功能且高性能的模型,甚至可部署于边缘设备进行推理。实验结果表明,该模型在多个基准测试中表现优异,为多模态人工智能的进一步发展铺平道路。
原文摘要 · Abstract (English)
We present a novel 4.5B parameter small language model that can handle multiple input and output modalities, including text, images, videos, and audio. Despite its small size, the model achieves near state-of-the-art performance on a variety of tasks, demonstrating the potential of multi-modal models to tackle complex real-world problems. Our approach leverages recent advancements in language modeling and multi-task learning to create a versatile and high-performing model that can even be deployed for edge inference. Experimental results show the model's strong performance across multiple benchmarks, paving the way for further progress in multi-modal artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。