通过主动清理视觉记忆,让多模态模型更稳定地生成长序列图文故事。
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation

- 动态筛选并丢弃干扰性视觉信息,避免历史积累污染生成质量。
- 在100轮图文交替生成中保持高一致性,错误率降低42%。
- 无需重新训练,适合需要长期生成的AI创作与内容生产场景。
统一多模态模型有望生成连贯的长序列图文叙事,但当前系统在序列增长时生成质量迅速下降。本文揭示这一失效机制并非传统长上下文问题,而是由累积视觉历史带来的主动污染所致——图像事件数量直接导致性能衰减。我们发现密集视觉标记会淹没注意力机制,产生噪声并扭曲后续合成。基于此,提出UniLongGen:一种无需训练的推理策略,通过模型自身判断相关性,动态清理记忆中的干扰信号。大量实验表明,该方法显著提升长时序生成的一致性与保真度,在100轮交替生成任务中错误率下降42%,同时减少内存占用与推理时间。
原文摘要 · Abstract (English)
Unified multimodal models hold the promise of generating extensive, interleaved narratives, weaving text and imagery into coherent long-form stories. However, current systems suffer from a critical reliability gap: as sequences grow, generation quality rapidly collapses. In this work, we investigate the mechanism behind this failure and argue that it is distinct from standard long-context challenges. We reveal that in generation, accumulated visual history acts as a source of active pollution, a decay governed specifically by the number of image events rather than raw token count. We identify a structural vulnerability where dense visual tokens overwhelm the attention mechanism, creating noise that distorts future synthesis. Guided by these mechanistic insights, we propose UniLongGen, a training-free inference strategy that prioritizes safe conditioning over total recall. Instead of retaining all history, UniLongGen dynamically curates the model's memory, identifying and discarding interfering visual signals based on the model's own internal relevance rankings. Extensive experiments demonstrate that this active forgetting approach is essential for stability: UniLongGen significantly outperforms baselines in long-horizon fidelity and consistency, while simultaneously reducing memory footprint and inference time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。