arXiv:2603.07697cs.CV2026-03中稿 · IEEE Transactions …被引 1

用自适应运动先验修复遮挡或噪声导致的运动数据缺失

Learning Context-Adaptive Motion Priors for Masked Motion Diffusion Models with Efficient Kinematic Attention Aggregation

  • 通过运动结构注意力聚合机制,高效编码关节与姿态特征
  • 在多种遮挡策略下实现优于现有方法的运动重建精度
  • 适合动作修复、补全和中间帧生成等场景,无需改动模型结构

基于视觉的动作捕捉常因遮挡导致关键关节信息丢失,影响三维动作重建精度;可穿戴设备则存在数据噪声或不稳定问题,需大量人工清洗。为此,我们提出掩码运动扩散模型(MMDM),一种基于扩散的生成重建框架,在掩码自编码器架构中利用部分高质量重构数据增强不完整或低置信度的运动数据。核心是运动学注意力聚合(KAA)机制,能高效、深度、迭代地编码关节级与姿态级特征,捕捉任务特定所需的结构与时间运动模式。本方法学习上下文自适应运动先验,由同一可复用架构提取的专用结构与时间特征,每种先验侧重不同运动动态,且对对应任务特别高效,实现无结构变更的自适应专业化。该特性使MMDM可高效适配动作修复、补全及中间帧生成等场景。在公开基准上的大量实验表明,其在多种掩码策略与任务设置下均表现优异。源代码已开源:https://github.com/jjkislele/MMDM。

原文摘要 · Abstract (English)

Vision-based motion capture solutions often struggle with occlusions, which result in the loss of critical joint information and hinder accurate 3D motion reconstruction. Other wearable alternatives also suffer from noisy or unstable data, often requiring extensive manual cleaning and correction to achieve reliable results. To address these challenges, we introduce the Masked Motion Diffusion Model (MMDM), a diffusion-based generative reconstruction framework that enhances incomplete or low-confidence motion data using partially available high-quality reconstructions within a Masked Autoencoder architecture. Central to our design is the Kinematic Attention Aggregation (KAA) mechanism, which enables efficient, deep, and iterative encoding of both joint-level and pose-level features, capturing structural and temporal motion patterns essential for task-specific reconstruction. We focus on learning context-adaptive motion priors, specialized structural and temporal features extracted by the same reusable architecture, where each learned prior emphasizes different aspects of motion dynamics and is specifically efficient for its corresponding task. This enables the architecture to adaptively specialize without altering its structure. Such versatility allows MMDM to efficiently learn motion priors tailored to scenarios such as motion refinement, completion, and in-betweening. Extensive evaluations on public benchmarks demonstrate that MMDM achieves strong performance across diverse masking strategies and task settings. The source code is available at https://github.com/jjkislele/MMDM.

动作补全扩散模型运动生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。