arXiv:2605.20777cs.CV2026-05中稿 · CVPR

让故事图像中的服饰细节更精准还原,提升视觉叙事的精细度。

AttriStory: Fine-grained Attribute Realization for Visual Storytelling with Diffusion Models

论文配图:AttriStory: Fine-grained Attribute Realization for Visual Storytelling with Diffusion Models
图 1 · 摘自论文原文
  • 通过早期去噪阶段的潜空间优化,精准定位服饰属性
  • 在10种艺术风格的200个多场景故事中实现属性一致
  • 可插拔集成,无需改动现有生成架构

基于扩散模型的视觉讲故事已显著提升角色跨场景的一致性,但对服装颜色、纹理等细粒度属性的忠实呈现仍缺乏系统方法。为此,我们提出AttriStory基准,利用大语言模型构建了涵盖10种艺术风格的200个多场景故事,每个场景均包含详细属性描述以支持丰富视觉叙事。针对属性实现问题,我们设计了一种可在早期去噪步骤中插入的即插即用潜空间优化模块,通过AttriLoss目标最大化期望属性-对象对的交叉注意力图对齐,同时抑制无关关联,引导模型准确定位属性。该方法与现有一致性机制正交,可无缝融入当前故事生成流程,无需架构修改。实验表明,加入AttriLoss后所有基线模型均获稳定提升。本工作将属性实现确立为与角色一致性并列的关键维度,推动细粒度可控叙事生成的发展。

原文摘要 · Abstract (English)

Visual storytelling with diffusion models has made impressive strides in maintaining character consistency across narrative scenes. However, a critical gap remains: while these methods ensure a character remains consistent across scenes, they provide no systematic method to ensure if fine-grained attributes such as color and textures of clothing, accessories are faithfully rendered in the generated images. Towards this goal, we introduce AttriStory, a benchmark enabling attribute realization in visual storytelling. We curate 200 multi-scene stories across 10 distinct artistic styles using Large Language Model. Each scene is constructed with detailed attribute specifications to enable rich visual narratives. Further, to address attribute realization, we propose a plug-and-play latent optimization module that operates during early denoising steps, when the model establishes structural and semantic content. We achieve this through AttriLoss objective designed to maximize alignment between the cross-attention maps for desired attribute-object pairs while suppressing spurious associations, guiding models to localize attributes correctly. This approach operates orthogonally to existing consistency mechanisms, integrating seamlessly with current story generation pipelines without requiring architectural modifications. Our experiments demonstrate consistent improvements on incorporating AttriLoss across all baselines. This work positions attribute realization as a distinct, complementary dimension of visual storytelling, alongside character consistency, advancing the field toward fine-grained attribute-controlled story generation. Project-page:https://manogna-s.github.io/attristory/

视觉叙事扩散模型属性控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。