arXiv:2503.14604cs.CVcs.AI2025-03IJCAI被引 25

MLLM时代图像描述评估面临新挑战,需改进评价方法。

Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives

  • 分析现有评估指标在多模态大模型下的表现
  • 发现传统指标对幻觉敏感且与人类判断相关性不足
  • 适合关注图像生成评估的科研人员参考

机器生成的图像描述评估是一项复杂且不断演进的挑战。随着多模态大语言模型(MLLMs)的出现,图像描述已成为核心任务,亟需可靠、稳健的评估指标。本文全面综述了图像描述评估领域的进展,分析了现有指标在演化过程中的优势与局限。我们从与人类判断的相关性、排名准确性及对幻觉的敏感性等多个维度评估指标性能。此外,探讨了MLLM生成更长、更详细的描述所带来的新挑战,并考察了现有指标对此类风格变化的适应能力。分析表明,标准评估方法存在局限,提出了未来研究的潜在方向。

原文摘要 · Abstract (English)

The evaluation of machine-generated image captions is a complex and evolving challenge. With the advent of Multimodal Large Language Models (MLLMs), image captioning has become a core task, increasing the need for robust and reliable evaluation metrics. This survey provides a comprehensive overview of advancements in image captioning evaluation, analyzing the evolution, strengths, and limitations of existing metrics. We assess these metrics across multiple dimensions, including correlation with human judgment, ranking accuracy, and sensitivity to hallucinations. Additionally, we explore the challenges posed by the longer and more detailed captions generated by MLLMs and examine the adaptability of current metrics to these stylistic variations. Our analysis highlights some limitations of standard evaluation approaches and suggests promising directions for future research in image captioning assessment.

图像描述评估方法多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。