arXiv:2608.31113cs.CV2026-08

用视频驱动3D模型动画,无需骨骼或绑定信息

BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives

论文配图:BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives
图 1 · 摘自论文原文
  • 用可学习的刚性运动基元组合生成动画,不依赖骨骼或权重
  • 在单目视频上实现高精度、时间稳定的3D网格动画
  • 适合需要快速生成动作的3D内容创作者

我们提出BLARM,一种用于视频驱动3D网格动画的前馈方法。给定单目视频和静态物体网格,BLARM预测一个时序连贯的动画网格,其运动与视频一致。不同于依赖显式骨架或直接回归高维顶点运动的方法,我们使用一组学习得到的、随时间变化的刚性运动基元和时不变的顶点-基元绑定权重来表示动画,从而在不需骨架、笼状结构、绑定权重或骨架标注的情况下,构建低维变形空间。架构通过分解的空间-时间注意力将几何衍生的变形隐变量与视频特征对齐,再通过预测的绑定权重混合刚性变换进行解码。训练采用轨迹重建、熵正则化和运动感知对比学习,使BLARM在单目视频上生成准确且时间稳定动画的同时,恢复出紧凑且可解释的运动结构。

原文摘要 · Abstract (English)

We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.

3D动画视频驱动运动基元单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。