构建多轮对话评估基准,测试视觉语言模型真实场景交互能力
MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- 从12个主流评测集整合647组多轮对话,每轮平均4次交互
- 18个模型在复杂对话中平均成功率仅50%,最强模型仍存短板
- 支持自动评估37项指标,适合研究多轮对话与上下文学习的学者
视觉语言模型(VLMs)在单轮任务中表现优异,但实际应用常需复杂多轮对话。现有数据集(如MMDU、ConvBench)未能充分覆盖真实对话场景。本文提出MultiVerse,一个包含647组对话的多轮对话基准,每组平均4轮,源自12个主流VLM评估集。涵盖484项任务和484个交互目标,涉及事实知识、感知理解到数学与编程等高级推理。为实现可靠评估,提出基于检查清单的自动化评测方法,利用GPT-4o评估37个关键维度,包括感知准确性、语言清晰度和事实正确性。在MultiVerse上评估18个VLM,结果显示即使最强模型(如GPT-4o)在复杂对话中成功率达50%,凸显其挑战性。值得注意的是,提供完整对话上下文显著提升小型或弱模型性能,强调了上下文学习的重要性。我们认为MultiVerse是评估VLM多轮交互能力的重要基准。
原文摘要 · Abstract (English)
Vision-and-Language Models (VLMs) have shown impressive capabilities on single-turn benchmarks, yet real-world applications often demand more intricate multi-turn dialogues. Existing multi-turn datasets (e.g, MMDU, ConvBench) only partially capture the breadth and depth of conversational scenarios encountered by users. In this work, we introduce MultiVerse, a novel multi-turn conversation benchmark featuring 647 dialogues - each averaging four turns - derived from a diverse set of 12 popular VLM evaluation benchmarks. With 484 tasks and 484 interaction goals, MultiVerse covers a wide range of topics, from factual knowledge and perception to advanced reasoning tasks such as mathematics and coding. To facilitate robust assessment, we propose a checklist-based evaluation method that leverages GPT-4o as the automated evaluator, measuring performance across 37 key aspects, including perceptual accuracy, linguistic clarity, and factual correctness. We evaluate 18 VLMs on MultiVerse, revealing that even the strongest models (e.g., GPT-4o) achieve only a 50% success rate in complex multi-turn conversations, highlighting the dataset's challenging nature. Notably, we find that providing full dialogue context significantly enhances performance for smaller or weaker models, emphasizing the importance of in-context learning. We believe MultiVerse is a landscape of evaluating multi-turn interaction abilities for VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。