arXiv:2605.13857cs.GRcs.CV2026-05被引 1

用扩散模型直接生成高保真动物动态视频,省去传统繁琐建模。

MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation

论文配图:MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation
图 1 · 摘自论文原文
  • 基于角色感知的旋转位置编码,实现动作同步与信息解耦。
  • 在120组动物网格-视频对上达成出色时序与结构一致性。
  • 适合影视特效、动画制作人员快速生成逼真动物动画。

电影级动物特效需精准模拟肌肉与毛发动态,但传统流程仍耗时且计算成本高昂。尽管生成式扩散模型在艺术创作中展现潜力,其在高保真动物模拟中的应用仍不充分。我们提出MoZoo,一种无需传统精修的生成式动力学求解器,可从粗略网格出发,在多模态引导下合成高质量动物视频。提出角色感知的旋转位置编码(RAR-RoPE),通过角色索引重映射实现运动对齐,并以固定时间偏移解耦参考信息。同时,采用非对称解耦注意力机制,将潜在序列分离以实现单向信息流,有效避免特征干扰并提升效率。针对高质量数据稀缺问题,构建了基于渲染引擎与逆映射的合成到真实数据管道,形成大规模成对序列数据集MoZoo-Data。进一步建立包含120组网格-视频对的MoZooBench基准。实验表明,MoZoo可在多种动物骨架与布局下实现高保真毛发模拟,保持优异的时序与结构一致性。

原文摘要 · Abstract (English)

The creation of cinematic-quality animal effects necessitates the precise modeling of muscle and fur dynamics, a process that remains both labor-intensive and computationally expensive within traditional production workflows. While generative diffusion models have shown promise in diverse artistic workflows, their capacity for high-fidelity animal simulation remains largely unexploited. We present MoZoo, a generative dynamics solver that bypasses conventional refinement to synthesize high-fidelity animal videos from coarse meshes under multimodal guidance. We propose Role-Aware RoPE (RAR-RoPE) which employs role-based index remapping to synchronize motion alignment while decoupling reference information via fixed temporal offsets. Complementing this, Asymmetric Decoupled Attention partitions the latent sequence to enforce a unidirectional information flow, effectively preventing feature interference and improving computational efficiency. To address the scarcity of high-quality training data, we introduce MoZoo-Data, a synthetic-to-real pipeline that leverages a rendering engine and an inverse mapping approach to construct a large-scale dataset of paired sequences. Furthermore, we establish MoZooBench, a comprehensive benchmark with 120 mesh-video pairs. Experimental results demonstrate that MoZoo achieves high-fidelity fur simulation across diverse animal skeletons and layouts, preserving superior temporal and structural consistency.

视频生成扩散模型动物模拟生成式建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。