让镜头与人体动作协同生成,拍出更符合电影逻辑的镜头画面。
Pulp Motion: Framing-aware multimodal camera and human motion generation
- 联合生成人体动作和镜头轨迹,通过屏幕构图实现双模态对齐。
- 在多个模型架构上提升构图一致性,文本对齐度也显著改善。
- 提出新数据集PulpMotion,适合影视动画、虚拟拍摄等场景研究者。
将人体动作与镜头轨迹分开生成会忽略电影创作的核心原则:演员表演与镜头运动在屏幕空间中的紧密互动。本文首次将该任务定义为文本条件下的联合生成,旨在保持一致的屏幕构图,同时生成两种异质但内在关联的模态:人体动作与镜头轨迹。我们提出一种简单且模型无关的框架,通过辅助模态——将人体关节投影到相机所形成的屏幕构图——来强制多模态一致性。这一构图作为模态间的自然桥梁,促进协调性并实现更精确的联合分布。我们首先设计了一个联合自编码器,学习共享潜在空间,并引入轻量级线性变换,从人体和镜头潜在表示映射到构图潜在空间。随后引入辅助采样机制,利用该线性变换引导生成过程朝向一致的构图模态。为支持此任务,我们还构建了PulpMotion数据集,包含丰富描述、高质量人体动作和镜头轨迹。在DiT与MAR-based架构上的大量实验表明,本方法在生成具有良好屏幕构图一致性的动作-镜头协同序列方面具有通用性和有效性,同时在双模态文本对齐方面均取得提升。定性结果显示更具电影逻辑的构图,达到该任务新基准。代码、模型与数据见项目主页。
原文摘要 · Abstract (English)
Treating human motion and camera trajectory generation separately overlooks a core principle of cinematography: the tight interplay between actor performance and camera work in the screen space. In this paper, we are the first to cast this task as a text-conditioned joint generation, aiming to maintain consistent on-screen framing while producing two heterogeneous, yet intrinsically linked, modalities: human motion and camera trajectories. We propose a simple, model-agnostic framework that enforces multimodal coherence via an auxiliary modality: the on-screen framing induced by projecting human joints onto the camera. This on-screen framing provides a natural and effective bridge between modalities, promoting consistency and leading to more precise joint distribution. We first design a joint autoencoder that learns a shared latent space, together with a lightweight linear transform from the human and camera latents to a framing latent. We then introduce auxiliary sampling, which exploits this linear transform to steer generation toward a coherent framing modality. To support this task, we also introduce the PulpMotion dataset, a human-motion and camera-trajectory dataset with rich captions, and high-quality human motions. Extensive experiments across DiT- and MAR-based architectures show the generality and effectiveness of our method in generating on-frame coherent human-camera motions, while also achieving gains on textual alignment for both modalities. Our qualitative results yield more cinematographically meaningful framings setting the new state of the art for this task. Code, models and data are available in our \href{https://www.lix.polytechnique.fr/vista/projects/2025_pulpmotion_courant/}{project page}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。