通过实体锚定调度生成顺序,提升多镜头视频视觉一致性。
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

- 基于实体首次出现时的视觉质量设定一致性上限,动态调度生成顺序。
- 在线构建实体视觉记忆库,生成前检索可靠参考以减少漂移。
- 无需训练或模型修改,适合需要高一致性的视频生成任务。
多镜头视频的视觉一致性仍是开放挑战。随着镜头数量增加,重复出现的实体(人物、物体、场景)易产生视觉漂移。我们发现观众判断一致性时,会将后续出现的实体与首次清晰呈现进行对比,初始视觉质量决定了后续一致性的上限。受此启发,提出无需训练、不依赖模型的代理式框架 GroundShot,通过在线构建实体级视觉记忆:根据镜头作为实体参考的预期价值调度生成顺序,从生成结果中提取并验证实体可靠性后存入记忆,生成新镜头前从记忆中检索合适参考。为评估该实体中心的一致性视角,进一步引入 GroundBench 基准,可分离控制性维度测量实体级一致性。实验表明,GroundShot 在无需额外训练或模型修改的前提下,显著优于现有方法。
原文摘要 · Abstract (English)
Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate across shots, causing entities that reappear across shots -- characters, objects, and locations -- to drift away from how they first appear. We observe that viewers judge consistency by comparing each later appearance of an entity with its first clear appearance; the visual quality of this initial appearance sets the consistency ceiling for all that follows. Motivated by this, we present \textbf{GroundShot}, a training-free, model-agnostic agentic framework for entity-grounded multi-shot generation. GroundShot builds an entity-level visual memory online from accepted generated shots: it schedules shots' generation order by their expected usefulness as entity references, grounds entities from generated videos, verifies their reliability before adding them to memory, and retrieves suitable entity references from memory before each shot is generated. To evaluate this entity-centered view of consistency, we further introduce \textbf{GroundBench}, a diagnostic benchmark that measures consistency at the entity level while isolating controlled challenge dimensions. Experiments show that GroundShot improves multi-shot consistency over existing methods while requiring no additional training or model modification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。