arXiv:2512.14420cs.CVcs.AI2025-12AAAI被引 2

DISCODE让图像描述评估更可靠,尤其在跨领域场景下接近人类判断。

DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning

  • 测试时自适应调整评分,用高斯先验优化分数估计。
  • 在六个领域的评估基准上表现超越现有方法。
  • 无需微调,适合希望提升评估鲁棒性的研究者。

大型视觉语言模型(LVLMs)在多种多模态任务中表现出色,但使用LVLM进行图像描述的鲁棒性自动评估仍具挑战性,尤其是在领域迁移场景下。为此,我们提出分布感知评分解码器(DISCODE),一种无需微调的新方法,可在不同领域间生成更贴近人类判断的稳健评估分数。其核心思想是测试时自适应评估,引入自适应测试时(ATT)损失,利用高斯先验分布提升评分估计的鲁棒性。该损失通过我们推导出的解析解在测试时高效最小化。此外,我们构建了多领域图像描述评估(MCEval)基准,涵盖六个不同领域,用于评估评估指标的鲁棒性。实验表明,DISCODE在MCEval及四个代表性基准上均达到参考无关评估指标的领先性能。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challenging, particularly under domain-shift scenarios. To address this issue, we introduce the Distribution-Aware Score Decoder (DISCODE), a novel finetuning-free method that generates robust evaluation scores better aligned with human judgments across diverse domains. The core idea behind DISCODE lies in its test-time adaptive evaluation approach, which introduces the Adaptive Test-Time (ATT) loss, leveraging a Gaussian prior distribution to improve robustness in evaluation score estimation. This loss is efficiently minimized at test time using an analytical solution that we derive. Furthermore, we introduce the Multi-domain Caption Evaluation (MCEval) benchmark, a new image captioning evaluation benchmark covering six distinct domains, designed to assess the robustness of evaluation metrics. In our experiments, we demonstrate that DISCODE achieves state-of-the-art performance as a reference-free evaluation metric across MCEval and four representative existing benchmarks.

图像描述自动评估鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。