FG-CLIP 2提升中英双语细粒度视觉语言对齐能力
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
- 融合区域-文本匹配与长描述建模,增强细粒度对齐
- 在29个数据集8项任务中均达领先水平,中文性能显著
- 支持中英文双语理解,适合多语言视觉推理研究者
细粒度视觉语言理解需要精准的视觉内容与语言描述对齐,现有模型在此方面仍受限,尤其在非英语场景。尽管CLIP等模型在全局对齐上表现良好,但难以捕捉物体属性、空间关系及语言表达的细微差别,且对双语理解支持不足。为此,我们提出FG-CLIP 2,一个专为中英文设计的细粒度视觉语言模型。该模型利用丰富的细粒度监督信号,包括区域-文本匹配和长描述建模,并引入文本内对比损失(TIC)以更好区分语义相近的描述。在精心构建的中英文大规模数据混合训练下,包括新发布的1200万条中文区域-文本数据集,FG-CLIP 2展现出强大的双语性能。为实现严谨评估,我们构建了新的中文多模态理解基准,涵盖长描述检索与边界框分类任务。在29个数据集、8项任务上的实验表明,该模型全面超越现有方法,达到当前最优水平。我们开源了模型、代码与评测基准,推动双语细粒度视觉语言对齐研究。
原文摘要 · Abstract (English)
Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP perform well on global alignment, they often struggle to capture fine-grained details in object attributes, spatial relations, and linguistic expressions, with limited support for bilingual comprehension. To address these challenges, we introduce FG-CLIP 2, a bilingual vision-language model designed to advance fine-grained alignment for both English and Chinese. Our approach leverages rich fine-grained supervision, including region-text matching and long-caption modeling, alongside multiple discriminative objectives. We further introduce the Textual Intra-modal Contrastive (TIC) loss to better distinguish semantically similar captions. Trained on a carefully curated mixture of large-scale English and Chinese data, including a newly released 12M Chinese region-text dataset, FG-CLIP 2 achieves powerful bilingual performance. To enable rigorous evaluation, we present a new benchmark for Chinese multimodal understanding, featuring long-caption retrieval and bounding box classification. Extensive experiments on 29 datasets across 8 tasks show that FG-CLIP 2 outperforms existing methods, achieving state-of-the-art results in both languages. We release the model, code, and benchmark to facilitate future research on bilingual fine-grained vision-language alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。