arXiv:2606.26171cs.CVcs.AI2026-06

解决长序列图像生成一致性问题,让漫画连贯不跑偏。

LCG: Long-Context Consistent Image Generation with Sparse Relational Attention

论文配图:LCG: Long-Context Consistent Image Generation with Sparse Relational Attention
图 1 · 摘自论文原文
  • 用稀疏关系注意力聚焦关键视觉特征,降低计算负担。
  • 引入路由一致性约束,多角色场景下也能保持形象一致。
  • 构建60万样本数据集,支持复杂叙事图像生成评测。

近期图像生成模型在单图合成上表现优异,但在漫画、分镜等连续输出任务中常出现一致性问题。本文提出长上下文生成(LCG)框架,通过稀疏关系注意力(SRA)机制,在扩展视觉上下文中选择性关注核心特征,确保语义与布局信息传播的可计算性。为强化语义对齐,设计路由一致性约束(RCC),利用身份感知掩码对齐生成分支间的结构模式,有效缓解复杂多角色场景中的外观漂移。为此构建了大规模合成数据集LCCD,包含60万条训练序列和1000条测试序列,每条序列含6至20张图像。实验表明,LCG在长上下文图像生成中显著优于基线模型,尤其在提示对齐与角色一致性方面表现更优。

原文摘要 · Abstract (English)

Recent image generation models achieve impressive quality in single-image synthesis, but often fail to maintain consistency across sequential outputs, as required in comics, storyboards, and visual narratives. We propose Long-Context Generation (LCG), a framework for long-context multi-image text-to-image generation, to improve consistency and scalability in long-context multi-image generation. LCG employs the Sparse Relational Attention (SRA) mechanism to selectively attend to core features across extended visual contexts, ensuring that the propagation of semantic and layout information remains computationally tractable. To enforce semantic alignment, we introduce the Routing Consistency Constraint (RCC), which leverages identity-aware masks to align structural patterns across generation branches, effectively mitigating drift in appearance even in complex multi-character scenes. To support training and evaluation in this setting, we construct the Long-Context Consistency Dataset (LCCD), a large-scale synthetic dataset comprising character-centric multi-image sequences spanning varied situational contexts. LCCD contains 600K training sequences and a separate 1K test set, with each sequence containing 6 to 20 images. The experiments demonstrate that LCG outperforms the compared baselines in prompt alignment and character consistency for long-context image generation, including multi-character scenes.

图像生成长序列一致性注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。