提升视觉对话与复杂指令跟随能力,解决多轮交互中的上下文丢失问题。
ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following
- 构建迭代式记忆-感知-规划-执行框架,增强视觉语言模型的持续理解力。
- 在300个复杂对话任务上平均得分4.03,优于GPT-4o和Gemini 1.5 Pro。
- 适合需要深度推理与多轮协作的多模态系统开发者与研究者。
尽管大语言模型(LLMs)和大视觉语言模型(LVLMs)取得显著进展,现有模型在处理复杂、多轮、视觉依存的任务时仍面临挑战,包括深层推理、持续上下文理解、实体追踪与多步指令遵循等问题。现有基准难以捕捉真实多模态交互的动态性与复杂性,导致上下文丢失与视觉幻觉。为此,我们提出MMDR-Bench(多模态对话推理基准),包含300个精心设计的复杂多轮对话场景,每轮平均5-7次交互,涵盖视觉实体追踪与推理深度等六项核心维度。同时,我们提出CoLVLM Agent(上下文型LVLM代理),通过迭代的“记忆-感知-规划-执行”循环,无需重新训练基础模型即可增强其推理与指令遵循能力。大量实验表明,CoLVLM Agent在MMDR-Bench上平均人类评分达4.03,显著优于GPT-4o(3.92)和Gemini 1.5 Pro(3.85),在推理深度、指令遵从性与错误抑制方面表现优异,且在长对话中保持稳定性能,验证了模块化设计与迭代方法的有效性。
原文摘要 · Abstract (English)
Despite significant advancements in Large Language Models (LLMs) and Large Vision-Language Models (LVLMs), current models still face substantial challenges in handling complex, multi-turn, and visually-grounded tasks that demand deep reasoning, sustained contextual understanding, entity tracking, and multi-step instruction following. Existing benchmarks often fall short in capturing the dynamism and intricacies of real-world multi-modal interactions, leading to issues such as context loss and visual hallucinations. To address these limitations, we introduce MMDR-Bench (Multi-Modal Dialogue Reasoning Benchmark), a novel dataset comprising 300 meticulously designed complex multi-turn dialogue scenarios, each averaging 5-7 turns and evaluated across six core dimensions including visual entity tracking and reasoning depth. Furthermore, we propose CoLVLM Agent (Contextual LVLM Agent), a holistic framework that enhances existing LVLMs with advanced reasoning and instruction following capabilities through an iterative "memory-perception-planning-execution" cycle, requiring no extensive re-training of the underlying models. Our extensive experiments on MMDR-Bench demonstrate that CoLVLM Agent consistently achieves superior performance, attaining an average human evaluation score of 4.03, notably surpassing state-of-the-art commercial models like GPT-4o (3.92) and Gemini 1.5 Pro (3.85). The framework exhibits significant advantages in reasoning depth, instruction adherence, and error suppression, and maintains robust performance over extended dialogue turns, validating the effectiveness of its modular design and iterative approach for complex multi-modal interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。