用自然语言直接生成视频中物体和镜头的3D运动轨迹
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
- 用大模型解析文本,生成符合影视语法的结构化运动程序
- 在真实数据集上实现比现有方法更精准的运动控制与意图对齐
- 适合需要复杂镜头调度的视频创作人员使用
视频生成在视觉质量和可控性方面取得显著进展,可基于文本、布局或运动进行条件生成。其中,运动控制——指定物体动态和相机轨迹——对构建复杂电影级场景至关重要,但现有交互方式仍受限。我们提出LAMP,利用大语言模型(LLMs)作为运动规划器,将自然语言描述转化为动态物体和相对定义的相机的显式3D轨迹。LAMP定义了一种受影视惯例启发的运动领域特定语言(DSL),借助LLMs的程序合成能力,从自然语言生成结构化的运动程序,并确定性地映射为3D轨迹。我们构建了一个大规模程序化数据集,包含自然文本描述与对应运动程序及3D轨迹的配对。实验表明,相比现有最优方法,LAMP在运动可控性和用户意图对齐方面表现更优,首次实现了从自然语言直接生成物体和相机运动的完整框架。代码、模型和数据已公开。
原文摘要 · Abstract (English)
Video generation has achieved remarkable progress in visual fidelity and controllability, enabling conditioning on text, layout, or motion. Among these, motion control - specifying object dynamics and camera trajectories - is essential for composing complex, cinematic scenes, yet existing interfaces remain limited. We introduce LAMP that leverages large language models (LLMs) as motion planners to translate natural language descriptions into explicit 3D trajectories for dynamic objects and (relatively defined) cameras. LAMP defines a motion domain-specific language (DSL), inspired by cinematography conventions. By harnessing program synthesis capabilities of LLMs, LAMP generates structured motion programs from natural language, which are deterministically mapped to 3D trajectories. We construct a large-scale procedural dataset pairing natural text descriptions with corresponding motion programs and 3D trajectories. Experiments demonstrate LAMP's improved performance in motion controllability and alignment with user intent compared to state-of-the-art alternatives establishing the first framework for generating both object and camera motions directly from natural language specifications. Code, models and data are available on our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。