arXiv:2608.08820cs.CV2026-08被引 1

让视频跨镜头逻辑连贯,避免内容断裂

LogiShot: Logically Coherent Cross-Shot Video Generation

论文配图:LogiShot: Logically Coherent Cross-Shot Video Generation
图 1 · 摘自论文原文
  • 用多模态线索联合编码上下文与条件信号
  • 通过视觉记忆保持跨镜头画面一致性
  • 首个专用于评估逻辑连贯性的数据集

生成具有逻辑连贯性的跨镜头视频对内容创作至关重要。当前多数跨镜头视频生成流程(如短剧制作)仍依赖孤立的文本脚本或明确参考图来指定生成内容。当用户指令不明确时,生成片段可能单个画面合理,但整体叙事断裂。为此,我们提出LogiShot,通过两条互补路径实现:1)联合编码上下文视频与其他条件信号,生成密集的多模态线索,提供跨镜头生成的视觉-语义依据;2)在生成过程中持续维护上下文视频的视觉记忆,以保持跨镜头视觉一致性。此外,我们构建了包含11万样本的数据集及专用评测基准。实验表明,LogiShot在多镜头逻辑连贯性上持续优于现有基线。模型与数据将公开可用。

原文摘要 · Abstract (English)

Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production, still rely on isolated textual scripts or explicit reference images to specify the generated content. Consequently, when user instructions are underspecified or ambiguous, a generated clip may appear visually plausible on its own but fail to align with the overall narrative, leading to disjointed content. We argue that achieving cross-shot logical coherence in video generation requires establishing logical connections across shots and maintaining visual consistency. To this end, we propose LogiShot, which incorporates information through two complementary paths: 1) LogiShot jointly encodes the context video and other conditioning signals, yielding dense multimodal cues that provide visual-semantic evidence for cross-shot generation; 2) the model maintains a visual memory of the context video throughout generation to preserve visual consistency across shots. Additionally, we construct a dataset with 110K samples and a dedicated benchmark for evaluating cross-shot logical coherence. Experiments demonstrate that LogiShot consistently outperforms existing baselines in terms of logical coherence across multiple shots. Model and data will be made publicly available.

视频生成逻辑连贯视觉记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。