提升多轮多模态对话的上下文理解能力,解决长对话记忆弱问题。
ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
- 引入记忆模块的上下文建模方法,增强多轮对话信息保留
- 在新构建的数据集上实现2%-4%的可用率提升
- 适合研究长对话、多模态交互的开发者和研究人员
多模态大模型展现出强大的零样本能力和图像理解性能,但现有开源模型在多轮交互尤其是长上下文场景下表现较弱。为此,本文提出一种新的上下文建模模块 ContextQFormer,通过引入记忆块增强上下文信息表达。同时,为促进后续研究,我们构建了一个全新的多轮多模态对话数据集 TMDialog,用于预训练、指令微调与评估,该数据集将后期开源。相比现有数据集,TMDialog包含更长的对话序列,支持多轮多模态对话研究。在 TMDialog 上与三个基线模型对比,实验结果表明,ContextQFormer 在可用率上相较基线提升 2%-4%。
原文摘要 · Abstract (English)
Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of multi-turn interaction, especially for long contexts. To address the issue, we first introduce a context modeling module, termed ContextQFormer, which utilizes a memory block to enhance the presentation of contextual information. Furthermore, to facilitate further research, we carefully build a new multi-turn multi-modal dialogue dataset (TMDialog) for pre-training, instruction-tuning, and evaluation, which will be open-sourced lately. Compared with other multi-modal dialogue datasets, TMDialog contains longer conversations, which supports the research of multi-turn multi-modal dialogue. In addition, ContextQFormer is compared with three baselines on TMDialog and experimental results illustrate that ContextQFormer achieves an improvement of 2%-4% in available rate over baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。