构建首个大规模多轮多图对话数据集,提升视觉语言模型的上下文推理能力。
MMCR: Advancing Visual Language Model in Multimodal Multi-Turn Contextual Reasoning
- 设计包含31万条多图多轮对话的数据集,每轮4~8次交互,覆盖1-4张图像。
- 在新基准上模型准确率提升5.2%,跨多个现有数据集也有1.1%-1.2%增益。
- 适合研究多模态对话、上下文理解与指令微调的学者和工程师使用。
相比单轮对话,涉及多张图片的多轮对话更贴近真实人机交互需求,且提供更丰富的上下文推理信息,有助于提升模型表现。然而,现有视觉语言模型主要依赖单轮对话训练与评估。本文基于人类对话特点(主题聚焦、内容简洁清晰),提出MMCR(Multimodal Multi-turn Contextual Reasoning)数据集:(1) MMCR-310k——目前最大的多图像多轮指令微调数据集,含31万条上下文对话,每条涵盖1-4张图像及4或8轮对话;(2) MMCR-Bench——诊断性基准,覆盖8个领域(人文、自然、科学、教育等)和40个子话题。大量实验表明,使用MMCR-310k微调的模型在MMCR-Bench上上下文准确率提升5.2%,并在现有基准上保持一致改进(AI2D+1.1%,MMMU+1.2%,MMVet+1.2%)。MMCR与提示工程将公开发布。
原文摘要 · Abstract (English)
Compared to single-turn dialogue, multi-turn dialogue involving multiple images better aligns with the needs of real-world human-AI interactions. Additionally, as training data, it provides richer contextual reasoning information, thereby guiding the model to achieve better performance. However, existing vision-language models (VLMs) primarily rely on single-turn dialogue training and evaluation benchmarks. In this paper, following the characteristics of human dialogue, such as focused topics and concise, clear content, we present MMCR (Multimodal Multi-turn Contextual Reasoning), a novel dataset comprising: (1) MMCR-310k -- the largest multi-image multi-turn instruction tuning dataset with 310K contextual dialogues, each covering 1-4 images and 4 or 8 dialogue turns; and (2) MMCR-Bench -- a diagnostic benchmark featuring dialogues, spanning 8 domains (Humanities, Natural, Science, Education, etc.) and 40 sub-topics. Extensive evaluations demonstrate that models fine-tuned with MMCR-310k achieve 5.2\% higher contextual accuracy on MMCR-Bench, while showing consistent improvements on existing benchmarks (+1.1\% on AI2D, +1.2\% on MMMU and MMVet). MMCR and prompt engineering will be released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。