用显式符号记忆解决图文交替生成中的逻辑漂移问题
Bringing The Consistency Gap: Explicit Structured Memory for Interleaved Image-Text Generation
- 构建图像理解树,以层级符号结构显式存储视觉信息
- 在3000样本测试中显著提升生成一致性,优于传统提示方法
- 适合需要长期上下文保持的复杂多模态交互场景
现有视觉语言模型在长时间、交错的图文交互中难以维持逻辑一致性、实体身份和艺术风格。我们将其归因于隐式神经表征在长序列中固有的衰减或纠缠问题,称为“多模态上下文漂移”。为此,提出IUT-Plug——一种与模型无关的神经符号结构状态追踪机制。该机制不依赖临时注意力图,而是引入图像理解树(IUT)作为显式持久记忆模块,通过解析视觉场景为层级符号结构(实体、属性、关系),增量更新并锁定不变属性、调整变化元素,并通过拓扑约束引导生成。我们在包含3000个人工标注样本的新基准上评估该方法,实验表明IUT-Plug能有效缓解上下文漂移,在一致性得分上显著优于无结构文本提示基线。结果验证了显式符号基础对多模态生成中长时一致性的重要性。
原文摘要 · Abstract (English)
Existing Vision Language Models (VLMs) often struggle to preserve logic, entity identity, and artistic style during extended, interleaved image-text interactions. We identify this limitation as "Multimodal Context Drift", which stems from the inherent tendency of implicit neural representations to decay or become entangled over long sequences. To bridge this gap, we propose IUT-Plug, a model-agnostic Neuro-Symbolic Structured State Tracking mechanism. Unlike purely neural approaches that rely on transient attention maps, IUT-Plug introduces the Image Understanding Tree (IUT) as an explicit, persistent memory module. The framework operates by (1) parsing visual scenes into hierarchical symbolic structures (entities, attributes, and relationships); (2) performing incremental state updates to logically lock invariant properties while modifying changing elements; and (3) guiding generation through topological constraints. We evaluate our approach on a novel benchmark comprising 3,000 human-annotated samples. Experimental results demonstrate that IUT-Plug effectively mitigates context drift, achieving significantly higher consistency scores compared to unstructured text-prompting baselines. This confirms that explicit symbolic grounding is essential for maintaining robust long-horizon consistency in multimodal generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。