融合第一视角视频与音乐,精准预测舞蹈动作
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
- 用骨架Mamba捕捉人体结构,建模动态运动
- 在36小时数据上超越现有方法,真实场景泛化强
- 适合舞蹈生成、虚拟人驱动等应用
人体舞蹈动作估计是具广泛工业应用的挑战性任务。近年来,研究多聚焦于仅使用第一视角视频或音乐输入进行动作预测,但联合利用两者进行动作估计仍鲜有探索。本文提出一种新方法,从第一视角视频和音乐中联合推断舞蹈动作。由于第一视角常遮挡身体,完整姿态估计困难;同时需使头部与躯干动作与视觉和音乐输入高度同步。我们构建了EgoAIST++数据集,包含超过36小时的跳舞动作,结合第一视角视图与音乐。受扩散模型与Mamba在序列建模中的成功启发,设计了以骨架Mamba为核心组件的EgoMusic Motion Network,显式建模人体骨骼结构。理论分析表明方法具有支持性。大量实验显示,本方法显著优于现有先进方法,并在真实数据上表现出良好泛化能力。
原文摘要 · Abstract (English)
Estimating human dance motion is a challenging task with various industrial applications. Recently, many efforts have focused on predicting human dance motion using either egocentric video or music as input. However, the task of jointly estimating human motion from both egocentric video and music remains largely unexplored. In this paper, we aim to develop a new method that predicts human dance motion from both egocentric video and music. In practice, the egocentric view often obscures much of the body, making accurate full-pose estimation challenging. Additionally, incorporating music requires the generated head and body movements to align well with both visual and musical inputs. We first introduce EgoAIST++, a new large-scale dataset that combines both egocentric views and music with more than 36 hours of dancing motion. Drawing on the success of diffusion models and Mamba on modeling sequences, we develop an EgoMusic Motion Network with a core Skeleton Mamba that explicitly captures the skeleton structure of the human body. We illustrate that our approach is theoretically supportive. Intensive experiments show that our method clearly outperforms state-of-the-art approaches and generalizes effectively to real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。