arXiv:2606.21135cs.CVcs.GR2026-06

首个融合身体形态的多模态动作生成框架,让不同体型的人动起来更真实。

Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion

论文配图:Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion
图 1 · 摘自论文原文
  • 基于生物形态信息生成动作,不把所有人当作相同体型处理。
  • 在文本、音乐、视频输入下均实现与体型匹配的动作输出,性能超越专用模型。
  • 支持无形态信息时同步恢复体型,统一估计与生成流程,适合动画/游戏应用。

人体动作生成已在文本、音乐、视频等多种模态下广泛研究,近期工作已将其统一于单一多模态框架中。然而,尽管性别和体型等形态因素会显著影响运动特征,现有统一框架仍忽略这些差异,将所有主体视为形态等同。我们提出 Odoriko,首个直接将受试者生物形态信息反映在生成动作中的统一多模态动作生成框架。不同于对个体差异进行平均,Odoriko 生成的动作与行动者身份一致,不仅反映任务指令,也体现其身体特征,涵盖文本、音乐、视频条件下的统一建模。当缺乏显式形态信息时,Odoriko 还能同步恢复主体形态,实现形态估计与动作生成的一体化。在文本到动作、音乐到舞蹈、视频到动作等多个基准上的大量实验表明,Odoriko 在标准指标上达到或超过现有专用模型水平,同时首次实现形态一致的动作生成,现有统一框架无法支持。

原文摘要 · Abstract (English)

Human motion generation has been widely studied across diverse input modalities, text, music, and video, and recent efforts have unified these into single multimodal frameworks. However, while morphological factors such as gender and body shape are known to produce distinct kinematic signatures, no existing unified framework incorporates this into generation, treating all subjects as morphologically equivalent. We present Odoriko, the first unified multimodal motion generation framework that reflects subject bio-morphological information directly in synthesized motion output. Rather than averaging over subject variation, Odoriko generates motion that is consistent with who is moving, not just what they are asked to do, across text, music, and video conditions within a single model. When explicit morphological information is unavailable, Odoriko additionally recovers subject morphology alongside motion, unifying estimation and generation in one framework. Extensive experiments across text-to-motion, music-to-dance, and video-to-motion benchmarks demonstrate that Odoriko matches or exceeds prior specialized models on standard metrics, while enabling morphology-consistent generation that no existing unified framework supports.

动作生成多模态扩散模型形态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。