arXiv:2605.06667cs.CVcs.AI2026-05International Conf…

零样本控制视频中角色动作与镜头运动,实现精准拍摄调度。

ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation

论文配图:ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
图 1 · 摘自论文原文
  • 通过分阶段条件引导,联合控制角色动作与镜头参数。
  • 在视角大幅变化下仍保持动作真实与镜头跟随准确。
  • 无需训练即可实现电影级运镜与动作同步,适合影视创作。

为艺术化视频生成需求,需对表演与镜头运动进行细粒度控制。本文提出ActCam,一种零样本视频生成方法,可将驱动视频中的角色动作迁移至新场景,并实现逐帧控制相机内在与外在参数。ActCam基于任意预训练的图像到视频扩散模型,接受场景深度和角色姿态作为条件。给定包含移动角色的源视频与目标相机运动,该方法生成跨帧几何一致的姿态与深度条件。随后进行单次采样,采用两阶段条件策略:早期去噪步骤同时依赖姿态与稀疏深度以维持场景结构;之后舍弃深度,仅用姿态引导优化高频细节,避免过度约束。在多个涵盖多样角色动作与复杂视角变换的基准上评估表明,相较于仅姿态控制及其他姿态与相机联合方法,ActCam显著提升镜头跟随精度与动作保真度,且在人类评估中更受青睐,尤其在大视角变化下表现优异。结果表明,精心设计的相机一致性条件与分阶段引导策略可在不训练的情况下实现强联合控制。

原文摘要 · Abstract (English)

For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two-phase conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose-only guidance refines high-frequency details without over-constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose-only control and other pose and camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations, especially under large viewpoint changes. Our results highlight that careful camera-consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/.

视频生成动作控制镜头调度扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。