构建多轮视觉对话数据集,提升模型跨模态理解与推理能力
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
- 设计规则+GPT辅助生成多轮对话数据,增强问题与图像关联性
- 提出DiagNote模型,通过双模块协同实现视觉-语言联合推理
- 在真实对话场景下显著优于现有模型,适合研究多轮对话的学者
多模态大语言模型(MLLM)在多模态理解方面表现出色,但多数模型仅基于单轮视觉问答任务训练,难以反映真实人类对话。本文提出MMDiag,一个协作生成的多轮多模态对话数据集,通过精心设计规则与GPT辅助,确保问题间、问题与图像间以及图像不同区域间的强相关性,更贴近真实场景。该数据集为多轮多模态对话学习提供强基准,对模型的定位与推理能力提出更高挑战。受人类视觉处理启发,我们提出DiagNote,一种具备多模态定位与推理能力的MLLM。其包含两个交互模块(沉思与凝视),分别执行思维链与标注,在多轮对话中协同工作。实验证明,相比现有MLLM,DiagNote在定位与联合处理视觉与语言信息方面具有显著优势。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs), built on large-scale pre-trained vision towers and language models, have shown great capabilities in multimodal understanding. However, most existing MLLMs are trained on single-turn vision question-answering tasks, which do not accurately reflect real-world human conversations. In this paper, we introduce MMDiag, a multi-turn multimodal dialogue dataset. This dataset is collaboratively generated through deliberately designed rules and GPT assistance, featuring strong correlations between questions, between questions and images, and among different image regions; thus aligning more closely with real-world scenarios. MMDiag serves as a strong benchmark for multi-turn multimodal dialogue learning and brings more challenges to the grounding and reasoning capabilities of MLLMs. Further, inspired by human vision processing, we present DiagNote, an MLLM equipped with multimodal grounding and reasoning capabilities. DiagNote consists of two modules (Deliberate and Gaze) interacting with each other to perform Chain-of-Thought and annotations respectively, throughout multi-turn dialogues. We empirically demonstrate the advantages of DiagNote in both grounding and jointly processing and reasoning with vision and language information over existing MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。