arXiv:2505.16180cs.CVcs.CL2025-05

用三重信号评估图像描述,更贴近人类判断。

Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation

  • 融合分布、感知和语言信号,综合评估图文匹配度。
  • 在Flickr8k上达58.42的Kendall-τ,优于多数现有方法。
  • 无需特定任务训练,适合跨数据集评测场景。

图像描述评估需兼顾视觉语义与语言语用,但多数指标难以全面捕捉。我们提出新型混合框架Redemption Score(RS),通过三角化三种互补信号进行排名:(1) 互信息发散(MID)衡量全局图文分布对齐;(2) 基于DINO的循环生成图像感知相似性,确保视觉锚定;(3) 大语言模型文本嵌入,对比上下文语义与人工参考。三者校准融合使RS提供更全面评估。在Flickr8k基准上,RS实现58.42的Kendall-τ,显著优于多数先前方法,并展现与人类判断更强的相关性,且无需任务特定训练。该框架在Conceptual Captions和MS COCO上也表现一致稳健,实现视觉准确性和文本质量的协同分析。

原文摘要 · Abstract (English)

Evaluating image captions requires cohesive assessment of both visual semantics and language pragmatics, which is often not entirely captured by most metrics. We introduce Redemption Score(RS), a novel hybrid framework that ranks image captions by triangulating three complementary signals: (1) Mutual Information Divergence (MID) for global image-text distributional alignment, (2) DINO-based perceptual similarity of cycle-generated images for visual grounding, and (3) LLM Text Embeddings for contextual text similarity against human references. A calibrated fusion of these signals allows RS to offer a more holistic assessment. On the Flickr8k benchmark, RS achieves a Kendall-$τ$ of 58.42, outperforming most prior methods and demonstrating superior correlation with human judgments without requiring task-specific training. Our framework provides a more robust and nuanced evaluation by thoroughly examining both the visual accuracy and text quality together, with consistent performance across Conceptual Captions and MS COCO.

图像描述评估框架多模态语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。