arXiv:2504.07556cs.CV2025-04CVPR被引 4

通过位置感知注意力提升图文细粒度对齐的评估精度

TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs

  • 用位置敏感的概率优化机制,聚焦关键语义词对应区域
  • 在NTIRE 2025挑战赛中获公开集第二(0.8445),私有集第二(0.8426)
  • 融合多模型视角的集成学习,适合评估生成模型的细节匹配能力

尽管近年来文本到图像生成模型取得了显著进展,但现有的视觉-语言对齐评估方法在细粒度语义匹配方面仍存在不足。基于全局相似度的现有方法常忽略文本与视觉内容之间的关键标记级对应关系。为此,我们提出TokenFocus-VQA,一种利用大视觉语言模型(LVLMs)的VQA范式进行位置特异性概率优化的新评估框架。核心创新在于设计了一种标记感知损失函数,可选择性地关注预定义词汇位置上的概率分布,从而精准衡量细粒度语义对齐。该框架进一步结合集成学习技术,从不同架构的LVLMs中聚合多视角评估结果,实现性能提升。在NTIRE 2025 T2I质量评估挑战赛第1赛道上,该方法在公开评估集排名第二(0.8445,仅比第一名低0.0001),在官方私有测试集上也位列第二(0.8426),展现出相比传统方法更强的捕捉细微图文对应关系的能力。

原文摘要 · Abstract (English)

While text-to-image (T2I) generation models have achieved remarkable progress in recent years, existing evaluation methodologies for vision-language alignment still struggle with the fine-grained semantic matching. Current approaches based on global similarity metrics often overlook critical token-level correspondences between textual descriptions and visual content. To this end, we present TokenFocus-VQA, a novel evaluation framework that leverages Large Vision-Language Models (LVLMs) through visual question answering (VQA) paradigm with position-specific probability optimization. Our key innovation lies in designing a token-aware loss function that selectively focuses on probability distributions at pre-defined vocabulary positions corresponding to crucial semantic elements, enabling precise measurement of fine-grained semantical alignment. The proposed framework further integrates ensemble learning techniques to aggregate multi-perspective assessments from diverse LVLMs architectures, thereby achieving further performance enhancement. Evaluated on the NTIRE 2025 T2I Quality Assessment Challenge Track 1, our TokenFocus-VQA ranks 2nd place (0.8445, only 0.0001 lower than the 1st method) on public evaluation and 2nd place (0.8426) on the official private test set, demonstrating superiority in capturing nuanced text-image correspondences compared to conventional evaluation methods.

图文对齐VQA评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。