测试多模态大模型在简单视觉指代任务中的实际理解能力
Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks?
- 用颜色块和色块网格测试模型指代理解能力
- 顶级模型仍难以准确解读依赖上下文的颜色描述
- 适合研究语言模型实用对话能力的学者参考
我们研究了多模态大语言模型在包含简单但抽象视觉刺激(如颜色块和色块网格)的指代消解任务中的语言能力。尽管该任务对当今的语言模型而言看似简单,且对人类双人对话者来说十分直接,但我们认为它是检验 MLLM 实用性能力的重要探针。我们的结果与分析表明,诸如基于上下文的颜色描述解读等基本语用能力,仍然是当前最先进 MLLM 的主要挑战。
原文摘要 · Abstract (English)
We investigate the linguistic abilities of multimodal large language models in reference resolution tasks featuring simple yet abstract visual stimuli, such as color patches and color grids. Although the task may not seem challenging for today's language models, being straightforward for human dyads, we consider it to be a highly relevant probe of the pragmatic capabilities of MLLMs. Our results and analyses indeed suggest that basic pragmatic capabilities, such as context-dependent interpretation of color descriptions, still constitute major challenges for state-of-the-art MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。