arXiv:2502.11300cs.CLcs.AI2025-02ACL被引 4

测试大模型理解图文连贯关系的能力,发现顶级模型表现不佳。

CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?

  • 构建新基准CORDIAL,评估多模态模型对图文连贯关系的理解。
  • 10+模型测试显示,顶级模型如Gemini 1.5 Pro仍低于简单分类器。
  • 呼吁用话语分析框架替代相似度指标,更精准评估模型能力。

多模态大语言模型(MLLMs)在指令遵循和跨领域推理方面表现出色,但现有评估基准主要关注事实与逻辑正确性,较少考察其对语用线索和跨模态关系的理解能力。为此,我们通过一致性关系(Coherence Relations)评估MLLMs在多模态话语分析(MDA)中的能力,提出新基准CORDIAL。该基准涵盖3个不同话语领域的多种一致性关系,具有不同粒度层级。在10+种不同提示策略下对10多个MLLM进行实验发现,即使顶尖模型如Gemini 1.5 Pro和GPT-4o,其性能也未能超越简单的分类器基线。研究强调应摒弃依赖相似性的评价方式,转而采用以话语驱动的评估框架,实现对模型能力更细致的刻画。基准与代码已公开:https://aashish2000.github.io/CORDIAL/

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are renowned for their superior instruction-following and reasoning capabilities across diverse problem domains. However, existing benchmarks primarily focus on assessing factual and logical correctness in downstream tasks, with limited emphasis on evaluating MLLMs' ability to interpret pragmatic cues and intermodal relationships. To address this gap, we assess the competency of MLLMs in performing Multimodal Discourse Analysis (MDA) using Coherence Relations. Our benchmark, CORDIAL, encompasses a broad spectrum of Coherence Relations across 3 different discourse domains at varying levels of granularity. Through our experiments on 10+ MLLMs employing different prompting strategies, we show that even top models like Gemini 1.5 Pro and GPT-4o fail to match the performance of simple classifier-based baselines. This study emphasizes the need to move beyond similarity-based metrics and adopt a discourse-driven framework for evaluating MLLMs, providing a more nuanced assessment of their capabilities. The benchmark and code are available at: https://aashish2000.github.io/CORDIAL/

多模态话语分析评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。