arXiv:2512.10352cs.CV2025-12

无需固定骨骼模板,可生成任意动物的文本驱动动作

Topology-Agnostic Animal Motion Generation from Text Prompt

  • 用统一框架建模任意骨骼结构与文本条件
  • 覆盖140种动物、3.3万段动作数据,支持跨物种迁移
  • 适合动画、虚拟角色和机器人动作生成研究者

动作生成是计算机动画的核心,广泛应用于娱乐、机器人和虚拟环境。现有方法多依赖固定骨骼模板,难以泛化到不同或变形的骨骼拓扑。本文针对当前方法缺乏大规模异构动物动作数据与统一生成框架的问题,提出OmniZoo——一个涵盖140个物种、32,979段序列的大型动物动作数据集,并附带多模态标注。基于此,我们构建了一种通用自回归动作生成框架,可为任意骨骼拓扑生成文本驱动的动作。核心是拓扑感知骨骼嵌入模块,将任意骨骼的几何与结构特性编码至共享令牌空间,实现与文本语义的无缝融合。给定文本提示与目标骨骼,该方法能生成时间连贯、物理合理且语义一致的动作,并支持跨物种动作风格迁移。

原文摘要 · Abstract (English)

Motion generation is fundamental to computer animation and widely used across entertainment, robotics, and virtual environments. While recent methods achieve impressive results, most rely on fixed skeletal templates, which prevent them from generalizing to skeletons with different or perturbed topologies. We address the core limitation of current motion generation methods - the combined lack of large-scale heterogeneous animal motion data and unified generative frameworks capable of jointly modeling arbitrary skeletal topologies and textual conditions. To this end, we introduce OmniZoo, a large-scale animal motion dataset spanning 140 species and 32,979 sequences, enriched with multimodal annotations. Building on OmniZoo, we propose a generalized autoregressive motion generation framework capable of producing text-driven motions for arbitrary skeletal topologies. Central to our model is a Topology-aware Skeleton Embedding Module that encodes geometric and structural properties of any skeleton into a shared token space, enabling seamless fusion with textual semantics. Given a text prompt and a target skeleton, our method generates temporally coherent, physically plausible, and semantically aligned motions, and further enables cross-species motion style transfer.

动作生成文本控制动物模拟拓扑不变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。