统一视觉与骨骼数据,实现动作感知与生成的端到端框架
Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation
- 设计视觉引导的动作分词器,联合学习视觉与骨骼信息
- 单模型处理视频动作估计、预测与补全,性能达最优或领先
- 适合需要跨模态动作理解与生成的研究者与开发者
人体动作分析任务如时序3D姿态估计、动作预测和动作插值在计算机视觉中至关重要。然而现有方法存在严重割裂:感知类模型从视频理解动作但仅输出文本,生成类模型无法直接处理原始视觉输入;现有生成式多模态大模型通常局限于单帧静态姿态,依赖密集参数化SMPL模型,难以处理时序动作;现有动作词汇库仅基于骨骼数据构建,脱离视觉域。为此,本文提出Superman,一个统一视觉感知与时序骨骼生成的框架。核心创新在于:首先,提出视觉引导的动作分词器,利用3D骨骼与视觉数据间的天然几何对齐,实现双模态鲁棒联合学习,构建统一的跨模态动作词汇;其次,在此动作语言基础上,训练单一统一的多模态大模型,灵活处理多样化时序输入,统一完成从视频中估计3D骨骼姿态(感知)以及基于骨骼的动作预测与插值(生成)。在Human3.6M等标准基准上的大量实验表明,该统一方法在所有动作任务上均达到当前最优或具有竞争力的性能,为基于骨骼的生成式动作分析提供更高效、可扩展的新路径。
原文摘要 · Abstract (English)
Human motion analysis tasks, such as temporal 3D pose estimation, motion prediction, and motion in-betweening, play an essential role in computer vision. However, current paradigms suffer from severe fragmentation. First, the field is split between ``perception'' models that understand motion from video but only output text, and ``generation'' models that cannot perceive from raw visual input. Second, generative MLLMs are often limited to single-frame, static poses using dense, parametric SMPL models, failing to handle temporal motion. Third, existing motion vocabularies are built from skeleton data alone, severing the link to the visual domain. To address these challenges, we introduce Superman, a unified framework that bridges visual perception with temporal, skeleton-based motion generation. Our solution is twofold. First, to overcome the modality disconnect, we propose a Vision-Guided Motion Tokenizer. Leveraging the natural geometric alignment between 3D skeletons and visual data, this module pioneers robust joint learning from both modalities, creating a unified, cross-modal motion vocabulary. Second, grounded in this motion language, a single, unified MLLM architecture is trained to handle all tasks. This module flexibly processes diverse, temporal inputs, unifying 3D skeleton pose estimation from video (perception) with skeleton-based motion prediction and in-betweening (generation). Extensive experiments on standard benchmarks, including Human3.6M, demonstrate that our unified method achieves state-of-the-art or competitive performance across all motion tasks. This showcases a more efficient and scalable path for generative motion analysis using skeletons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。