零样本生成连贯多角色动作叙事,保持身份与背景一致
StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives
- 三模块协同:动态聚焦主体、强化动作语义、智能保留背景线索
- CLIP-T提升10-15%,背景一致性优于基线,24GB显存下推理更快
- 适合需要多角色连续视觉叙事的创作场景,无需微调
生成无需微调的多帧、高动作密度视觉叙事面临三大挑战:动作文本忠实度、主体身份保真度与跨帧背景连续性。我们提出StoryTailor,一个零样本流水线,可在单张RTX 4090(24 GB)上运行,从长篇叙述提示、每主体参考图像及定位框生成时序连贯、身份稳定的图像序列。系统由三个协同模块驱动:基于高斯中心的注意力(GCA)动态聚焦主体核心,缓解定位框重叠;动作增强奇异值重加权(AB-SVR)放大文本嵌入空间中的动作相关方向;选择性遗忘缓存(SFC)保留可迁移背景线索,遗忘非必要历史,并选择性输出保留线索以建立跨场景语义关联。实验表明,相比基线方法,CLIP-T最高提升10-15%,DreamSim低于强基线,而CLIP-I保持在视觉可接受且有竞争力的范围内。在相同分辨率与步数下,推理速度优于FluxKontext。定性结果显示,StoryTailor能生成富有表现力的互动与持续演变但稳定的场景。
原文摘要 · Abstract (English)
Generating multi-frame, action-rich visual narratives without fine-tuning faces a threefold tension: action text faithfulness, subject identity fidelity, and cross-frame background continuity. We propose StoryTailor, a zero-shot pipeline that runs on a single RTX 4090 (24 GB) and produces temporally coherent, identity-preserving image sequences from a long narrative prompt, per-subject references, and grounding boxes. Three synergistic modules drive the system: Gaussian-Centered Attention (GCA) to dynamically focus on each subject core and ease grounding-box overlaps; Action-Boost Singular Value Reweighting (AB-SVR) to amplify action-related directions in the text embedding space; and Selective Forgetting Cache (SFC) that retains transferable background cues, forgets nonessential history, and selectively surfaces retained cues to build cross-scene semantic ties. Compared with baseline methods, experiments show that CLIP-T improves by up to 10-15%, with DreamSim lower than strong baselines, while CLIP-I stays in a visually acceptable, competitive range. With matched resolution and steps on a 24 GB GPU, inference is faster than FluxKontext. Qualitatively, StoryTailor delivers expressive interactions and evolving yet stable scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。