让视频生成支持多种输入方式控制镜头运动。
TriMotion: Modality-Agnostic Camera Control for Video Generation

- 统一视频、姿态、文本输入的镜头运动表示空间。
- 在三种输入下均实现高精度镜头轨迹跟随。
- 适合需要灵活镜头控制的视频生成应用。
镜头运动控制对生成式系统中的视角变化至关重要。然而,现有方法通常依赖单一模态输入(如显式姿态轨迹或参考视频),难以支持多样化的用户输入。为此,我们提出TriMotion,一个模态无关的相机控制视频生成框架,可将描述相同镜头运动的视频、姿态和文本输入映射到共享的运动嵌入空间。为学习该空间,我们构建了Motion Triplet Dataset,通过扩展多相机视频数据集并基于相机外参生成几何精确的运动描述。此外,我们引入潜在运动一致性目标,利用运动嵌入空间直接在潜在空间中引导生成视频遵循目标镜头轨迹,避免像素空间解码开销。大量实验表明,TriMotion在三种模态下均能生成高质量且准确跟随目标镜头轨迹的视频。除标准生成外,共享运动嵌入空间还支持序列化运动组合与跨模态运动插值等灵活应用。
原文摘要 · Abstract (English)
Camera motion control is essential for directing viewpoint changes in generative systems. However, existing methods typically condition the generation process on a single specific modality, such as explicit pose trajectories or reference videos, limiting their ability to support heterogeneous user inputs. To address this limitation, we present TriMotion, a modality-agnostic framework for camera-controlled video generation that maps video, pose, and text inputs, describing the same camera trajectory into a shared motion embedding space. Learning such a space requires synchronized supervision across modalities. Therefore, we build the Motion Triplet Dataset by extending a Multi-Cam Video Dataset with geometry-grounded motion descriptions derived from camera extrinsics. We further introduce a latent motion consistency objective that leverages the motion embedding space to encourage the generated video to follow the target camera trajectory directly in latent space, avoiding the cost of pixel-space decoding. Extensive experiments show that TriMotion generates high-quality videos that accurately follow the target camera trajectories across all three modalities. Beyond standard generation, the shared motion embedding space also enables flexible applications such as sequential motion composition and cross-modal motion interpolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。