arXiv:2411.09691cs.CV2024-11EMNLP被引 14

通过多尺度对齐提升视觉细节理解能力

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

  • 引入多尺度图文对齐机制,融合文本、坐标与图像信息
  • 构建超30万条合成数据集,显著增强细粒度对齐效果
  • 轻量级模型仅30亿参数,性能媲美大型多模态模型

多模态大语言模型在细粒度视觉理解任务中表现卓越,但常因缺乏细粒度知识对齐而难以准确捕捉局部细节并实现全面全局感知。现有方法虽关注对象描述与定位信息的对齐,却通常未显式整合包含丰富语义的图像信息。为此,本文提出一种新颖的细粒度视觉知识对齐方法,有效融合包括文本、坐标和图像在内的多尺度对象知识。该方法依托于自研的多尺度细粒度增强数据合成流水线,生成超过30万条高质量训练数据,显著提升对齐精度与整体性能。此外,我们提出一系列轻量级模型TinyGroundingGPT,参数量约30亿,在复杂视觉场景下表现出色,其定位性能可媲美更大规模的多模态模型。

原文摘要 · Abstract (English)

Multi-modal large language models (MLLMs) have achieved remarkable success in fine-grained visual understanding across a range of tasks. However, they often encounter significant challenges due to inadequate alignment for fine-grained knowledge, which restricts their ability to accurately capture local details and attain a comprehensive global perception. While recent advancements have focused on aligning object expressions with grounding information, they typically lack explicit integration of object images, which contain affluent information beyond mere texts or coordinates. To bridge this gap, we introduce a novel fine-grained visual knowledge alignment method that effectively aligns and integrates multi-scale knowledge of objects, including texts, coordinates, and images. This innovative method is underpinned by our multi-scale fine-grained enhancement data synthesis pipeline, which provides over 300K essential training data to enhance alignment and improve overall performance. Furthermore, we present TinyGroundingGPT, a series of compact models optimized for high-level alignments. With a scale of approximately 3B parameters, TinyGroundingGPT achieves outstanding results in grounding tasks while delivering performance comparable to larger MLLMs in complex visual scenarios.

多模态细粒度理解视觉对齐轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。