arXiv:2607.26910cs.CV2026-07

用大模型生成符合镜头语言的3D摄像机运动,让文字描述变电影级视频。

CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents

论文配图:CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents
图 1 · 摘自论文原文
  • 用大模型理解文本,拆解镜头动作为推轨、环绕等原子动作
  • 基于3D场景图生成无碰撞、高画质的摄像机轨迹,提升镜头表达力
  • 适合影视创作、虚拟导览等需要自然语言生成视频的场景

从自然语言描述中自动生成具有电影表现力的3D场景摄像机轨迹是一项极具实用价值的任务,广泛应用于房地产广告、虚拟导览等领域。现有方法或依赖2D图像先验而缺乏真正的3D空间感知,或把轨迹生成当作脱离电影语义的几何路径规划问题。我们提出CinemaTraj,将摄像机轨迹规划重构为语言引导的空间推理任务。给定一组RGB-D图像和用户提示,CinemaTraj为大模型代理构建结构化的3D场景图:代理将提示分解为一系列原子镜头动作(推轨、环绕、升降、平移、俯仰、变焦、弧线)。每个动作通过一种新型参数化轨迹表示实现,兼具电影表现力与可优化的防碰撞特性。场景图作为结构化空间先验,使代理的推理基于环境的精确几何与语义知识。CinemaTraj还同步生成与镜头运动对齐的语音旁白和字幕,输出带解说的电影级视频。我们在真实世界的ScanNet++环境中评估,结果表明其生成的轨迹忠实于提示、无碰撞且具备高电影质量,在提示对齐度、轨迹质量和安全性上均优于现有方法。

原文摘要 · Abstract (English)

Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high practical value, with applications ranging from real-estate advertising to virtual tour creation. Existing methods either lack true 3D spatial awareness by relying on 2D image priors, or treat trajectory generation as a geometric path planning problem divorced from cinematographic semantics. We present CinemaTraj, a framework that reframes camera trajectory planning as a language-grounded spatial reasoning problem. Given a set of RGB-D images and a user prompt, CinemaTraj equips an LLM agent with a structured 3D scene graph: the agent decomposes the prompt into a sequence of atomic cinematographic movements (dolly, orbit, crane, pan, tilt, zoom, arc). Each movement is instantiated via a novel parametric trajectory representation that is both cinematographically expressive and optimizable for collision avoidance. The scene graph acts as a structured spatial prior, grounding the agent's reasoning in accurate geometric and semantic knowledge of the environment. CinemaTraj further generates synchronized voiceover and subtitles aligned with camera motion, producing narrated cinematic video outputs. We evaluate CinemaTraj on real-world ScanNet++ environments, and show that it produces prompt-faithful, collision-free trajectories with high cinematographic quality, outperforming existing approaches on prompt alignment, trajectory quality, and safety metrics.

视频生成大模型3D轨迹电影语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。