用提示词层记忆图修复长剧生成中的视觉不连贯问题
SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale

- 通过提取每镜头多维状态构建记忆图,仅回溯因果相关上下文
- 在跨集连续性召回率上从0.700提升至0.946,提升24.6个百分点
- 无需训练、兼容主流模型,适合大规模视频生成生产流程
短剧生成已发展为工业化流程,但由单镜头孤立生成导致视觉连续性瓶颈。现有框架各镜头独立生成,引发角色姿态、布景、道具等属性漂移,累积后形成严重视觉断裂。本文提出SEAM(Shot Entity-Attribute Memory),一种无需训练、模型无关的记忆图机制,通过提取每镜头的多维状态,在提示词层重构上下文:仅检索因果前序信息,选择性过滤,并以自然语言重写方式注入约束。我们进一步发布SEAM-Bench双盲连续性分镜基准,实验显示,SEAM将跨集连续性召回率从0.700提升至0.946,泛化于六种主流文本模型,并在生成图像层实现稳定但未达显著的改进。部署于CreativeFitting的SEAM-Agent生产流水线中,处理201个镜头时达到96.5%导演通过率,无安全风险注入;保守反事实分析表明,跨集记忆贡献至少21.9个百分点的通过率提升。
原文摘要 · Abstract (English)
Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual continuity has become a critical bottleneck. Current agent frameworks generate each shot in isolation, so context drifts across shots and props, character posture, and blocking turn inconsistent. Once assembled, these small discrepancies amplify into severe visual breaks. We present SEAM (Shot Entity-Attribute Memory), a training-free, model-agnostic memory graph that repairs continuity entirely at the prompt-text layer by extracting a multi-dimensional state for every shot, retrieving only causally prior context over the resulting graph, filtering it selectively, and injecting the surviving constraints by natural-language prompt rewriting. We further release SEAM-Bench, a double-blind continuity storyboarding benchmark, on which SEAM raises cross-episode continuity recall from 0.700 to 0.946, generalizes across six mainstream text models, and yields consistent, though not yet significant, gains at the generated-image layer. Deployed as a mandatory stage in CreativeFitting's SEAM-Agent production pipeline over 201 shots, SEAM reaches a 96.5% director-acceptance rate with zero unsafe injections; a conservative counterfactual attributes at least 21.9 percentage points of that rate to its cross-episode memory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。