一个模型同时搞定视觉和语言任务,性能超越现有水平。
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
- 构建多维度数据集CCMD-8M,统一视觉与语言任务训练
- 提出Griffon-G模型,在多个任务上达专家级表现
- 适合需要通用多模态能力的研究者和开发者
大型多模态模型(LMMs)在视觉-语言和视觉中心任务中取得显著进展,但多数模型仅专注其中一类。本文提出全新多维度数据集CCMD-8M,通过多层次数据清洗与多任务整合,突破两类任务融合的数据障碍。在此基础上,提出Griffon-G——一个端到端统一处理视觉-语言与视觉中心任务的通用大模型。该模型解决联合优化中的训练崩溃问题,提升训练效率。在多模态基准、通用视觉问答(VQA)、场景文本相关VQA、文档类VQA、指代理解及目标检测等任务上的评估显示,Griffon-G超越先进LMMs,在复杂视觉中心任务中达到专家级性能。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have achieved significant breakthroughs in various vision-language and vision-centric tasks based on auto-regressive modeling. However, these models typically focus on either vision-centric tasks, such as visual grounding and region description, or vision-language tasks, like image caption and multi-scenario VQAs. None of the LMMs have yet comprehensively unified both types of tasks within a single model, as seen in Large Language Models in the natural language processing field. Furthermore, even with abundant multi-task instruction-following data, directly stacking these data for universal capabilities extension remains challenging. To address these issues, we introduce a novel multi-dimension curated and consolidated multimodal dataset, named CCMD-8M, which overcomes the data barriers of unifying vision-centric and vision-language tasks through multi-level data curation and multi-task consolidation. More importantly, we present Griffon-G, a general large multimodal model that addresses both vision-centric and vision-language tasks within a single end-to-end paradigm. Griffon-G resolves the training collapse issue encountered during the joint optimization of these tasks, achieving better training efficiency. Evaluations across multimodal benchmarks, general Visual Question Answering (VQA) tasks, scene text-centric VQA tasks, document-related VQA tasks, Referring Expression Comprehension, and object detection demonstrate that Griffon-G surpasses the advanced LMMs and achieves expert-level performance in complicated vision-centric tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。