arXiv:2606.21661cs.CV2026-06被引 3

用记忆门控实现多镜头音视频连贯生成,支持跨镜头身份一致性。

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

论文配图:UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating
图 1 · 摘自论文原文
  • 设计双记忆槽+边界感知门控,固定大小存储长期与短期视觉记忆。
  • 在多镜头生成中超越开源基线,媲美闭源最强系统。
  • 适合需要跨镜头一致性控制的影视生成与多语言内容创作。

生成连贯的多镜头视频需结构化跨镜头记忆。人物外观、场景上下文与说话人身份必须在镜头切换间保持一致。现有方法或在固定长度序列上端到端训练难以扩展,或逐镜头生成时记忆库线性增长,或依赖大语言模型规划但缺乏多镜头感知的骨干网络。我们提出 UnityShots,基于 LTX-2.3 模型,使用标注的电影与音乐视频镜头数据进行训练。视频流维护两个固定大小的记忆槽:锚定首镜头的长期记忆(LTM)和保存前一镜头尾部的短期记忆(STM),每个切镜处由融合视觉切镜概率与节拍追踪信号的边界条件门控更新。音频流在每镜头注入参考说话人令牌以保留声纹,无需滑动音频记忆库。通过 AdaLN 学习的离散切镜类型先验,在推理时作为控制过渡强度的调节旋钮。我们发布一个包含200个跨文化多镜头序列的基准数据集,涵盖六个民族区域与十种以上语言,提供每镜头参考身份、参考音频及边界过渡标签。在 I2V、T2V、R2V 条件模式下评估,UnityShots 在所有跨镜头连贯性指标上均优于开源基线,并在多镜头性能上达到最强闭源系统水平。

原文摘要 · Abstract (English)

Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches either train end-to-end over fixed-length sequences and cannot scale, generate shot-by-shot with memory banks that grow linearly, or orchestrate pretrained generators under an LLM planner without a multi-shot-aware backbone. We present UnityShots, a memory-driven multi-shot audio-video generation system built on LTX-2.3, trained on annotated cinematic and music-video shots. The video stream maintains two fixed-size slots, a long-term memory (LTM) slot anchored to the opening shot and a short-term memory (STM) slot holding the immediately preceding tail, both updated at every cut by a boundary-conditioned gate that fuses visual cut probability and beat-tracker signals. The audio stream injects a reference speaker token at every shot to preserve vocal timbre without a sliding audio bank. A discrete cut-type prior, learned through AdaLN, becomes an inference-time control knob over transition strength. We release a benchmark of $200$ multi-cultural multi-shot sequences spanning six ethnic regions and ten or more languages, with per-shot reference identities, reference audio, and per-boundary transition labels. Evaluated across I2V, T2V, and R2V conditioning modes, UnityShots leads open-source baselines on every cross-shot coherence metric and matches the strongest closed-source system on the multi-shot axes.

多镜头生成记忆机制音视频同步跨镜头一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。