arXiv:2603.11421cs.CV2026-03被引 4

用文本生成电影镜头序列,自动规划运镜轨迹。

ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation

  • 分两阶段:先用视觉语言模型规划全局运镜路径,再由控制器生成视频。
  • 在多镜头场景下实现运镜精准、跨镜头连贯,效果优于现有方法。
  • 构建了高质量影视数据集,适合影视创作与自动化视频生成研究者。

文本驱动的视频生成已降低电影创作门槛,但电影级多镜头场景中的摄像机控制仍是关键瓶颈。隐式文本提示精度不足,显式轨迹输入则需大量手动操作且易失败。为此,我们提出以数据为中心的新范式,认为(字幕、轨迹、视频)三元组存在内在联合分布,可连接自动规划与精确执行。基于此,我们设计ShotVerse——一种“先规划后控制”框架,由视觉语言模型驱动的规划器利用空间先验从文本生成具有全局一致性的电影级运镜轨迹,控制器通过摄像机适配器将轨迹转化为多镜头视频内容。核心在于构建数据基础:设计自动化多镜头摄像机校准流程,将分散的单镜头轨迹对齐至统一全局坐标系,从而形成ShotVerse-Bench数据集,支持三轨评估。大量实验表明,ShotVerse有效弥合了不可靠文本控制与高成本人工规划之间的鸿沟,在摄像机准确性与跨镜头一致性上表现优异。

原文摘要 · Abstract (English)

Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Implicit textual prompts lack precision, while explicit trajectory conditioning imposes prohibitive manual overhead and often triggers execution failures in current models. To overcome this bottleneck, we propose a data-centric paradigm shift, positing that aligned (Caption, Trajectory, Video) triplets form an inherent joint distribution that can connect automated plotting and precise execution. Guided by this insight, we present ShotVerse, a ``Plan-then-Control'' framework that decouples generation into two collaborative agents: a VLM (Vision-Language Model)-based Planner that leverages spatial priors to obtain cinematic, globally aligned trajectories from text, and a Controller that renders these trajectories into multi-shot video content via a camera adapter. Central to our approach is the construction of a data foundation: we design an automated multi-shot camera calibration pipeline aligns disjoint single-shot trajectories into a unified global coordinate system. This facilitates the curation of ShotVerse-Bench, a high-fidelity cinematic dataset with a three-track evaluation protocol that serves as the bedrock for our framework. Extensive experiments demonstrate that ShotVerse effectively bridges the gap between unreliable textual control and labor-intensive manual plotting, achieving superior cinematic aesthetics and generating multi-shot videos that are both camera-accurate and cross-shot consistent.

视频生成运镜控制多镜头文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。