arXiv:2605.13838cs.CVcs.GR2026-05International Conf…

解决3D动画中模型初始姿态与视频不匹配的问题,实现精准动态生成。

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

论文配图:R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow
图 1 · 摘自论文原文
  • 通过分解输入为基底网格、运动轨迹和修正偏移量,自动对齐初始姿态。
  • 在50万+序列的数据集上验证,显著减少几何畸变,提升动画成功率。
  • 适合需要高精度3D角色动画的创作者和工业级内容生成场景。

视频引导的3D动画在内容创作中潜力巨大,但实际应用面临一个常被忽视的关键挑战:姿态错位困境。用户提供的静态网格初始姿态通常与参考视频首帧不一致,直接强制跟随会导致严重几何畸变或动画失败。为此,我们提出统一框架R-DMesh,生成与视频上下文对齐的高保真4D网格。不同于常规运动迁移方法,本方法引入新型变分自编码器(VAE),显式解耦输入为条件基底网格、相对运动轨迹和关键的修正跳跃偏移量。该偏移量被学习以自动将输入网格的任意姿态调整至与视频初始状态对齐。我们通过三流注意力机制处理这些组件,利用顶点级几何特征调制三个正交流动,确保修正与动画过程中的物理一致性与局部刚性。生成阶段采用基于修正流的扩散变换器,以预训练视频潜在表示为条件,有效将丰富的时空先验迁移到3D域。为支持该任务,我们构建了包含超过50万条动态网格序列的Video-RDMesh数据集,专门模拟姿态错位问题。大量实验表明,R-DMesh不仅解决了对齐难题,还支持鲁棒的下游应用,如姿态重定向与全链路4D生成。

原文摘要 · Abstract (English)

Video-guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets. However, practical deployment faces a critical yet frequently overlooked hurdle: the pose misalignment dilemma. In real-world scenarios, the initial pose of a user-provided static mesh rarely aligns with the starting frame of a reference video. Naively forcing a mesh to follow a mismatched trajectory inevitably leads to severe geometric distortion or animation failure. To address this, we present Rectified Dynamic Mesh (R-DMesh), a unified framework designed to generate high-fidelity 4D meshes that are ``rectified'' to align with video context. Unlike standard motion transfer approaches, our method introduces a novel VAE that explicitly disentangles the input into a conditional base mesh, relative motion trajectories, and a crucial rectification jump offset. This offset is learned to automatically transform the arbitrary pose of the input mesh to match the video's initial state before animation begins. We process these components via a Triflow Attention mechanism, which leverages vertex-wise geometric features to modulate the three orthogonal flows, ensuring physical consistency and local rigidity during the rectification and animation process. For generation, we employ a Rectified Flow-based Diffusion Transformer conditioned on pre-trained video latents, effectively transferring rich spatio-temporal priors to the 3D domain. To support this task, we construct Video-RDMesh, a large-scale dataset of over 500k dynamic mesh sequences specifically curated to simulate pose misalignment. Extensive experiments demonstrate that R-DMesh not only solves the alignment problem but also enables robust downstream applications, including pose retargeting and holistic 4D generation.

3D动画视频生成姿态对齐扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。