arXiv:2603.15614cs.CV2026-03被引 1

统一控制视频场景、主体和动作,实现更精准的AI视频生成。

Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion

论文配图:Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
图 1 · 摘自论文原文
  • 通过三重提示框架整合场景、主体和运动控制
  • 在多视角一致性与动作准确性上超越现有模型
  • 适合需要精细视频编辑的创作者或影视制作人员

近期视频扩散模型在视觉质量上取得显著进展,但精确的细粒度控制仍是限制内容创作实用性的关键瓶颈。对AI视频创作者而言,三种控制至关重要:(i) 场景构图,(ii) 多视角一致的主体定制,(iii) 相机姿态或物体运动调节。现有方法通常孤立处理这些维度,对多视角主体合成和任意姿态变化下的身份保持支持有限。为此,我们提出Tri-Prompting,一种统一框架与两阶段训练范式,整合场景构图、多视角主体一致性与运动控制。该方法利用由3D追踪点驱动的双条件运动模块处理背景场景,用下采样RGB线索处理前景主体。为平衡可控性与视觉真实感,进一步提出推理阶段的ControlNet缩放调度。Tri-Prompting支持新工作流,包括将主体3D感知地插入任意场景,以及对图像中已有主体进行操作。实验表明,其在多视角主体身份、3D一致性与运动精度上显著优于Phantom和DaS等专用基线模型。

原文摘要 · Abstract (English)

Recent video diffusion models have made remarkable strides in visual quality, yet precise, fine-grained control remains a key bottleneck that limits practical customizability for content creation. For AI video creators, three forms of control are crucial: (i) scene composition, (ii) multi-view consistent subject customization, and (iii) camera-pose or object-motion adjustment. Existing methods typically handle these dimensions in isolation, with limited support for multi-view subject synthesis and identity preservation under arbitrary pose changes. This lack of a unified architecture makes it difficult to support versatile, jointly controllable video. We introduce Tri-Prompting, a unified framework and two-stage training paradigm that integrates scene composition, multi-view subject consistency, and motion control. Our approach leverages a dual-condition motion module driven by 3D tracking points for background scenes and downsampled RGB cues for foreground subjects. To ensure a balance between controllability and visual realism, we further propose an inference ControlNet scale schedule. Tri-Prompting supports novel workflows, including 3D-aware subject insertion into any scenes and manipulation of existing subjects in an image. Experimental results demonstrate that Tri-Prompting significantly outperforms specialized baselines such as Phantom and DaS in multi-view subject identity, 3D consistency, and motion accuracy.

视频生成扩散模型3D控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。