iMatch通过多策略增强,精准评估图文生成的语义对齐。
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
- 用指令增强微调多模态大模型,提升图文匹配评估精度。
- 在CVPR NTIRE 2025竞赛中排名第一,显著优于现有方法。
- 适合关注生成质量评估、图文对齐研究的研究者使用。
随着文本到图像(T2I)生成模型的快速发展,评估生成图像与文本描述之间的语义对齐成为重要挑战。现有基于视觉问答(VQA)的方法仍难以实现细粒度评估和精确量化。本文提出一种名为iMatch的改进评估方法,通过微调多模态大语言模型来评估图像-文本语义对齐。提出四种创新增强策略:1)QAlign将模型离散评分转换为连续匹配分数;2)利用模型预测的伪标签扩充验证集,提升泛化能力;3)引入元素类别标签,增强对图像-文本匹配的理解;4)采用随机光照等图像增强技术,提高模型鲁棒性。此外,还设计提示类型增强和分数扰动策略以进一步提升元素评估准确性。实验表明,iMatch显著优于现有方法,并在CVPR NTIRE 2025文本到图像生成模型质量评估-第一赛道中获得第一名。
原文摘要 · Abstract (English)
With the rapid advancement of text-to-image (T2I) generation models, assessing the semantic alignment between generated images and text descriptions has become a significant research challenge. Current methods, including those based on Visual Question Answering (VQA), still struggle with fine-grained assessments and precise quantification of image-text alignment. This paper presents an improved evaluation method named Instruction-augmented Multimodal Alignment for Image-Text and Element Matching (iMatch), which evaluates image-text semantic alignment by fine-tuning multimodal large language models. We introduce four innovative augmentation strategies: First, the QAlign strategy creates a precise probabilistic mapping to convert discrete scores from multimodal large language models into continuous matching scores. Second, a validation set augmentation strategy uses pseudo-labels from model predictions to expand training data, boosting the model's generalization performance. Third, an element augmentation strategy integrates element category labels to refine the model's understanding of image-text matching. Fourth, an image augmentation strategy employs techniques like random lighting to increase the model's robustness. Additionally, we propose prompt type augmentation and score perturbation strategies to further enhance the accuracy of element assessments. Our experimental results show that the iMatch method significantly surpasses existing methods, confirming its effectiveness and practical value. Furthermore, our iMatch won first place in the CVPR NTIRE 2025 Text to Image Generation Model Quality Assessment - Track 1 Image-Text Alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。