arXiv:2604.25164cs.CV2026-04被引 1

让动作生成模型学会根据身体特征调整动作,更真实

IAM: Identity-Aware Human Motion and Shape Joint Generation

论文配图:IAM: Identity-Aware Human Motion and Shape Joint Generation
图 1 · 摘自论文原文
  • 用语言和图像隐含表征身份,不依赖具体体型数据
  • 同时生成动作和身体形状,提升动作与身份一致性
  • 在真人视频和动捕数据上验证,动作更自然合理

当前文本驱动的人体动作生成方法大多假设身份中性,使用标准身体表示生成动作,忽略了身体形态对运动动态的显著影响。实际上,身体比例、质量分布和年龄等属性会显著影响动作表现,忽略这种耦合常导致物理不一致的动作。我们提出一种身份感知的动作生成框架,显式建模身体形态与运动动态之间的关系。身份通过自然语言描述和视觉线索等多模态信号隐式表示,无需显式几何测量。我们进一步引入动作-形状联合生成范式,同步合成动作序列与身体形状参数,使身份信息直接调控运动动态。在动捕数据集和大规模野外视频上的实验表明,该方法在保持高动作质量的同时,显著提升了动作的真实感与动作-身份一致性。

原文摘要 · Abstract (English)

Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches assume identity-neutral motion and generate movements using a canonical body representation, ignoring the strong influence of body morphology on motion dynamics. In practice, attributes such as body proportions, mass distribution, and age significantly affect how actions are performed, and neglecting this coupling often leads to physically inconsistent motions. We propose an identity-aware motion generation framework that explicitly models the relationship between body morphology and motion dynamics. Instead of relying on explicit geometric measurements, identity is represented using multimodal signals, including natural language descriptions and visual cues. We further introduce a joint motion-shape generation paradigm that simultaneously synthesizes motion sequences and body shape parameters, allowing identity cues to directly modulate motion dynamics. Extensive experiments on motion capture datasets and large-scale in-the-wild videos demonstrate improved motion realism and motion-identity consistency while maintaining high motion quality. Project page: https://vjwq.github.io/IAM

动作生成身份感知联合生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。