arXiv:2606.02000cs.CVcs.AI2026-06

用3D网格令牌实现无需渲染的人体动作控制,让视频扩散模型真正理解三维结构。

Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization

论文配图:Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization
图 1 · 摘自论文原文
  • 用压缩的3D人体网格令牌替代2D渲染图像作为生成条件
  • 在人体动作控制任务上表现优异,减少视角依赖和轨迹-姿态错配问题
  • 适合需要精确3D人体建模与动作编辑的研究者

扩散模型在视频生成中表现出色,但其是否真正理解视觉观测背后的三维结构,而非仅生成合理的二维投影,仍是一个未解问题。本文通过人体动作控制任务来探究该问题,该任务需要精确建模三维人体几何、运动、相机视角和场景上下文。不同于依赖渲染2D动作引导视频的先前方法,本文提出一种无需渲染的框架,直接以压缩的3D人体网格令牌作为视频生成的条件。该表示保留完整三维几何信息,同时支持统一的令牌化生成流程,在基于DiT的架构中联合处理视频令牌与动作令牌。这一设计要求模型在生成过程中同时推理外观、三维结构与相机视角。实验结果表明,该方法在人体动作控制基准测试中表现良好,同时减少了由视点依赖的2D引导和编辑过程中的轨迹-姿态不匹配所导致的伪影。这些发现表明,配备网格令牌化的视频扩散模型能够更好地捕捉复杂的人体三维结构及其与周围环境的交互。

原文摘要 · Abstract (English)

Diffusion models have shown remarkable success in video generation. However, whether such models are truly aware of the 3D structure underlying visual observations, rather than simply reproducing plausible 2D projections, remains an open question. In this work, we investigate this question through human motion control, a task that requires precise modelling of 3D human geometry, motion, camera viewpoint, and scene context. Unlike prior methods that rely on rendered 2D motion guidance videos, we propose a render-free framework that conditions video generation directly on compressed 3D human mesh tokens. This representation preserves full 3D geometric information while enabling a unified token-based generation pipeline that processes video tokens jointly with motion tokens in a DiT-based architecture. This design requires the model to reason jointly about appearance, 3D structure, and camera viewpoint during video generation. Experimental results demonstrate strong performance on human motion control benchmarks, while reducing artifacts induced by view-dependent 2D guidance and trajectory-pose mismatches during editing. These findings suggest that video diffusion models, when equipped with mesh tokenization, can better capture complex 3D human structures and their interactions with the surrounding environment.

3D感知视频生成扩散模型人体动作控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。