arXiv:2602.21273cs.CV2026-02中稿 · CVPR

零样本生成连贯多角色动作叙事,保持身份与背景一致

StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives

  • 三模块协同:动态聚焦主体、强化动作语义、智能保留背景线索
  • CLIP-T提升10-15%,背景一致性优于基线,24GB显存下推理更快
  • 适合需要多角色连续视觉叙事的创作场景,无需微调

生成无需微调的多帧、高动作密度视觉叙事面临三大挑战:动作文本忠实度、主体身份保真度与跨帧背景连续性。我们提出StoryTailor,一个零样本流水线,可在单张RTX 4090(24 GB)上运行,从长篇叙述提示、每主体参考图像及定位框生成时序连贯、身份稳定的图像序列。系统由三个协同模块驱动:基于高斯中心的注意力(GCA)动态聚焦主体核心,缓解定位框重叠;动作增强奇异值重加权(AB-SVR)放大文本嵌入空间中的动作相关方向;选择性遗忘缓存(SFC)保留可迁移背景线索,遗忘非必要历史,并选择性输出保留线索以建立跨场景语义关联。实验表明,相比基线方法,CLIP-T最高提升10-15%,DreamSim低于强基线,而CLIP-I保持在视觉可接受且有竞争力的范围内。在相同分辨率与步数下,推理速度优于FluxKontext。定性结果显示,StoryTailor能生成富有表现力的互动与持续演变但稳定的场景。

原文摘要 · Abstract (English)

Generating multi-frame, action-rich visual narratives without fine-tuning faces a threefold tension: action text faithfulness, subject identity fidelity, and cross-frame background continuity. We propose StoryTailor, a zero-shot pipeline that runs on a single RTX 4090 (24 GB) and produces temporally coherent, identity-preserving image sequences from a long narrative prompt, per-subject references, and grounding boxes. Three synergistic modules drive the system: Gaussian-Centered Attention (GCA) to dynamically focus on each subject core and ease grounding-box overlaps; Action-Boost Singular Value Reweighting (AB-SVR) to amplify action-related directions in the text embedding space; and Selective Forgetting Cache (SFC) that retains transferable background cues, forgets nonessential history, and selectively surfaces retained cues to build cross-scene semantic ties. Compared with baseline methods, experiments show that CLIP-T improves by up to 10-15%, with DreamSim lower than strong baselines, while CLIP-I stays in a visually acceptable, competitive range. With matched resolution and steps on a 24 GB GPU, inference is faster than FluxKontext. Qualitatively, StoryTailor delivers expressive interactions and evolving yet stable scenes.

视觉叙事零样本多角色生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。