精准生成双人交互动作,考虑距离与层级动态变化。
Fine-grained text-driven dual-human motion generation via dynamic hierarchical interaction
- 分三阶段建模:个体文本分解、动态距离预测、整体引导精修。
- 在双人动作数据集上优于现有方法,动作更精细自然。
- 适合需要真实双人交互生成的动画、游戏开发场景。
人类互动具有内在的动态性与层次性:动态指动作随距离变化,层次则从个体到人际互动再到整体动作。现有方法多忽略距离与层次,采用静态建模。为此,本文提出细粒度双人动作生成方法 FineDual,通过三阶段建模动态层次化交互。第一阶段(自学习阶段)利用大语言模型将双人总体文本拆分为个体文本,对齐个体级文本与动作特征;第二阶段(自适应调整阶段)通过交互距离预测器估计交互距离,借助交互感知图网络在人际层面动态建模交互;第三阶段(教师引导精修阶段)以整体文本特征为指导,在整体层面优化动作特征,生成细粒度高质量双人动作。在双人动作数据集上的大量定量与定性评估表明,FineDual 有效建模了动态层次化人体交互,显著优于现有方法。
原文摘要 · Abstract (English)
Human interaction is inherently dynamic and hierarchical, where the dynamic refers to the motion changes with distance, and the hierarchy is from individual to inter-individual and ultimately to overall motion. Exploiting these properties is vital for dual-human motion generation, while existing methods almost model human interaction temporally invariantly, ignoring distance and hierarchy. To address it, we propose a fine-grained dual-human motion generation method, namely FineDual, a tri-stage method to model the dynamic hierarchical interaction from individual to inter-individual. The first stage, Self-Learning Stage, divides the dual-human overall text into individual texts through a Large Language Model, aligning text features and motion features at the individual level. The second stage, Adaptive Adjustment Stage, predicts interaction distance by an interaction distance predictor, modeling human interactions dynamically at the inter-individual level by an interaction-aware graph network. The last stage, Teacher-Guided Refinement Stage, utilizes overall text features as guidance to refine motion features at the overall level, generating fine-grained and high-quality dual-human motion. Extensive quantitative and qualitative evaluations on dual-human motion datasets demonstrate that our proposed FineDual outperforms existing approaches, effectively modeling dynamic hierarchical human interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。