arXiv:2412.14613cs.CLcs.AI2024-12被引 2

用视觉语言模型实现多任务多标准自动评估,更贴近人工判断。

Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models

  • 自底向上融合各评价标准得分,生成综合评分
  • 在18,000条人类标注上验证,相关性优于传统指标
  • 适合需要多维度评估多模态生成质量的研究者

视觉语言模型(VLMs)在多种多模态任务中表现出色。然而,现有用于评估VLM生成文本质量的指标通常仅针对特定任务进行整体评价,如图像描述。尽管整体评价至关重要,但不同任务关注的标准各异,导致现有指标难以适应多任务场景。为此,我们提出HarmonicEval——一种无参考的综合性评估指标,通过自底向上聚合各准则得分来生成总体分数。为进一步评估自动评估指标在多任务场景中的泛化能力,我们构建了多任务多标准人类评估基准MMHE,包含4个跨模态任务下的18,000条专家人工判断。实验表明,HarmonicEval在与人工判断的相关性上优于传统指标,同时提供每项准则的具体数值评分。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have shown impressive abilities across a range of multi-modal tasks. However, existing metrics for evaluating the quality of text generated by VLMs typically focus on an overall evaluation for a specific task, such as image captioning. While the overall evaluation is essential for any task, the criteria prioritized can differ depending on the task, making it challenging for current metrics to adapt to multi-task scenarios. To address this limitation, we propose HarmonicEval, a reference-free comprehensive evaluation metric that aggregates criterion-wise scores to produce the overall score in a bottom-up manner. Furthermore, to assess the generalizability of automatic evaluation metrics in multi-task scenarios, we construct the Multi-task Multi-criteria Human Evaluation (MMHE) benchmark, which comprises 18,000 expert human judgments across four multi-modal tasks. Our experiments demonstrate that HarmonicEval achieves higher correlations with human judgments than conventional metrics while providing numerical scores for each criterion. Project page: https://stjohn2007.github.io/MMHE_project/

多模态评估视觉语言模型自动评估人类判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。