用稀疏3D框实现视频场景的直观导演式控制。
LooseControlVideo: Directorial Video Control using Spatial Blocking

- 用稀疏定向3D框作为布局代理,替代密集标注。
- 轨迹误差降低1.2至3倍,遮挡准确率提升1.5至2倍。
- 适合需要精细空间控制的多对象动态视频创作。
文本生成视频中的精确3D空间编排仍是重大挑战,尤其在多物体场景中,语义布局与时间动态常纠缠不清。现有深度条件模型虽结构保真度高,但需密集、逐帧精准引导,对动态变形物体难以高效制作。我们提出LooseControlVideo,通过稀疏定向3D框作为“遮挡”代理,实现直观且富有表现力的控制,用户可设定高层布局与轨迹,由视频生成模型自动补全真实遮挡、运动与交互。方法基于在带有DNOCS标注的视频数据集上微调Wan 2.2骨干网络,DNOCS是一种新型3D尺寸、朝向及深度排序遮挡编码。此外,支持局部优化,如调整跳跃轨迹或添加交互,对全局场景影响极小。在nuScenes、HO-3D和BEHAVE基准测试中,相比现有2D框与流基基线,显著提升:轨迹误差降低1.2至3倍;刚体运动一致性提升2倍;遮挡准确率提升1.5至2倍,证明定向3D原语为复杂多智能体视频创作提供了良好几何先验。
原文摘要 · Abstract (English)
Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and temporal dynamics are often entangled. While existing depth-conditioned models achieve good structural fidelity, they necessitate dense, frame-accurate guidance that is labor-intensive to author for dynamic events involving deformable objects. We present LooseControlVideo, a framework that enables intuitive and expressive control by using sparse, oriented 3D boxes as a "blocking" proxy. This allows users to author high-level layout and trajectory while leveraging a video generative model to generate realistic occlusions, dynamics and interactions. We achieve this by fine-tuning a Wan 2.2 backbone on a video dataset annotated with DNOCS, a novel encoding for 3D size, orientation and depth-ordered occlusions. Furthermore, our method allows for localized refinement, such as adjusting a jump trajectory or adding an interaction, with minimal disruption to the global scene context. Extensive evaluations on the nuScenes, HO-3D, and BEHAVE benchmarks demonstrate that LooseControlVideo significantly outperforms existing 2D-box and flow-based baselines. Our findings indicate a 1.2x to 3x improvement in Trajectory Error; 2x improvement in Rigid Motion Consistency; and a 1.5x to 2x increase in Occlusion Accuracy over current state-of-the-art layout-conditioned models, demonstrating that oriented 3D primitives provide good geometric prior for complex, multi-agent video authoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。