让视觉语言模型学会连贯多轮推理,避免遗忘和幻觉。
Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance
- 用动态记忆单元存储跨轮次的图文信息
- 多轮推理准确率超越现有方法,尤其在长对话中表现更优
- 适合需要深度上下文理解的对话系统研究者
当前大语言模型和视觉语言大模型在单轮任务中表现优异,但在需要深层上下文理解与复杂视觉推理的多轮交互中,常出现推理碎片化、上下文丢失和幻觉问题。为此,我们提出上下文感知多轮视觉-文本推理框架CAMVR,通过两个核心创新:视觉-文本上下文记忆单元(VCMU),一种可动态读写、存储关键视觉特征、文本语义表示及其跨模态对应关系的记忆网络;以及自适应视觉焦点引导机制(AVFG),利用VCMU中的上下文动态调整视觉编码器对图像相关区域的关注。多层次推理融合策略确保生成回答与当前输入及累积历史上下文高度一致。在VisDial、改造版A-OKVQA及新构建的多轮指令遵循(MTIF)数据集上的大量实验表明,CAMVR持续达到领先性能。
原文摘要 · Abstract (English)
Current Large Language Models (LLMs) and Vision-Language Large Models (LVLMs) excel in single-turn tasks but face significant challenges in multi-turn interactions requiring deep contextual understanding and complex visual reasoning, often leading to fragmented reasoning, context loss, and hallucinations. To address these limitations, we propose Context-Aware Multi-Turn Visual Reasoning (CAMVR), a novel framework designed to empower LVLMs with robust and coherent multi-turn visual-textual inference capabilities. CAMVR introduces two key innovations: a Visual-Textual Context Memory Unit (VCMU), a dynamic read-write memory network that stores and manages critical visual features, textual semantic representations, and their cross-modal correspondences from each interaction turn; and an Adaptive Visual Focus Guidance (AVFG) mechanism, which leverages the VCMU's context to dynamically adjust the visual encoder's attention to contextually relevant image regions. Our multi-level reasoning integration strategy ensures that response generation is deeply coherent with both current inputs and accumulated historical context. Extensive experiments on challenging datasets, including VisDial, an adapted A-OKVQA, and our novel Multi-Turn Instruction Following (MTIF) dataset, demonstrate that CAMVR consistently achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。