构建真实对话场景的多模态大模型评测基准,揭示其记忆与推理缺陷。
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation
- 构建包含5120次真实对话的多模态对话评测集,覆盖六大核心能力
- 20个模型在持续对话中准确率下降,暴露记忆衰退等四大缺陷
- 提出记笔记策略,显著提升多轮对话表现,适合对话系统研究者
近期多模态大语言模型(MLLMs)在开放性对话中展现出显著潜力,能生成更准确、个性化的回应。然而,它们在真实场景下长期交互中的记忆、回忆和推理能力仍缺乏深入探索。本文提出MMRC——一个多模态真实对话基准,用于评估MLLMs的六项核心开放性能力:信息提取、多轮推理、信息更新、图像管理、记忆召回和拒绝回答。数据来自真实场景,包含5,120次对话和28,720个手动标注问题,对现有MLLMs构成重大挑战。在20个MLLM上的评估显示,开放性对话中准确率下降。我们识别出四种常见失败模式:长期记忆退化、事实知识更新不足、错误累积传播、不愿说不。为此,提出简单有效的记笔记策略,可记录对话关键信息并在回复时提醒模型,显著提升对话能力。跨六个MLLM的实验验证了该策略的有效性。
原文摘要 · Abstract (English)
Recent multimodal large language models (MLLMs) have demonstrated significant potential in open-ended conversation, generating more accurate and personalized responses. However, their abilities to memorize, recall, and reason in sustained interactions within real-world scenarios remain underexplored. This paper introduces MMRC, a Multi-Modal Real-world Conversation benchmark for evaluating six core open-ended abilities of MLLMs: information extraction, multi-turn reasoning, information update, image management, memory recall, and answer refusal. With data collected from real-world scenarios, MMRC comprises 5,120 conversations and 28,720 corresponding manually labeled questions, posing a significant challenge to existing MLLMs. Evaluations on 20 MLLMs in MMRC indicate an accuracy drop during open-ended interactions. We identify four common failure patterns: long-term memory degradation, inadequacies in updating factual knowledge, accumulated assumption of error propagation, and reluctance to say no. To mitigate these issues, we propose a simple yet effective NOTE-TAKING strategy, which can record key information from the conversation and remind the model during its responses, enhancing conversational capabilities. Experiments across six MLLMs demonstrate significant performance improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。