arXiv:2412.05983cs.CV2024-12ICCV被引 11

让通用模型学会专业领域知识,用专家模型提升特定任务表现

Chimera: Improving Generalist Model with Domain-Specific Experts

  • 通过渐进式训练融合专家模型特征到通用模型输入
  • 在图表、表格、数学等场景实现顶尖多模态推理性能
  • 适合需要跨领域视觉理解的开发者与研究者

近期大型多模态模型(LMMs)的发展表明,增加图文配对数据规模可显著提升通用任务表现。然而,这类通用模型主要基于以自然图像为主的网络规模数据集训练,导致在需大量领域先验知识的特定任务上能力不足。同时,直接整合针对特定领域的专家模型面临表征差异和优化不平衡的问题。为此,我们提出Chimera,一种可扩展且低成本的多模态流程,用于增强现有LMMs的领域专用能力。具体而言,设计渐进式训练策略,将专家模型特征融入通用模型输入;为缓解因通用视觉编码器高度对齐带来的优化不平衡问题,引入新型通用-专家协同掩码(GSCM)机制。该方法使模型在图表、表格、数学及文档等多领域表现出色,在多模态推理与视觉内容提取任务上达到当前最优水平,这些任务是评估现有LMMs的关键挑战。

原文摘要 · Abstract (English)

Recent advancements in Large Multi-modal Models (LMMs) underscore the importance of scaling by increasing image-text paired data, achieving impressive performance on general tasks. Despite their effectiveness in broad applications, generalist models are primarily trained on web-scale datasets dominated by natural images, resulting in the sacrifice of specialized capabilities for domain-specific tasks that require extensive domain prior knowledge. Moreover, directly integrating expert models tailored for specific domains is challenging due to the representational gap and imbalanced optimization between the generalist model and experts. To address these challenges, we introduce Chimera, a scalable and low-cost multi-modal pipeline designed to boost the ability of existing LMMs with domain-specific experts. Specifically, we design a progressive training strategy to integrate features from expert models into the input of a generalist LMM. To address the imbalanced optimization caused by the well-aligned general visual encoder, we introduce a novel Generalist-Specialist Collaboration Masking (GSCM) mechanism. This results in a versatile model that excels across the chart, table, math, and document domains, achieving state-of-the-art performance on multi-modal reasoning and visual content extraction tasks, both of which are challenging tasks for assessing existing LMMs.

多模态专家模型领域适应视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。