arXiv:2602.19004cs.CV2026-02中稿 · CVPR被引 1

让惯性传感器与视频姿态精准对齐,实现秒级动作同步定位。

MoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment

  • 用骨骼运动替代原始图像,聚焦真实动作信号
  • 分部位对齐传感器与姿态,实现亚秒级时间匹配
  • 适合需要高精度动作同步的可穿戴设备研究

本文旨在学习惯性测量单元(IMU)信号与视频提取的2D姿态序列之间的联合表示,以实现跨模态检索、时间同步、主体与身体部位定位及动作识别。为此,提出MoBind,一种分层对比学习框架,解决三个挑战:(1)过滤无关视觉背景,(2)建模多传感器IMU结构配置,(3)实现细粒度、亚秒级时间对齐。为提取运动相关特征,MoBind将IMU信号与骨骼运动序列对齐,而非原始像素。进一步将全身运动分解为局部身体部位轨迹,每个部位对应其对应的IMU,实现语义对齐的多传感器对齐。为捕捉精细时间对应关系,采用分层对比策略:先对齐帧级时间片段,再融合局部(部位)对齐与全局(全身)运动聚合。在mRi、TotalCapture和EgoHumans数据集上评估,MoBind在所有四项任务中均显著优于强基线,展现出鲁棒的细粒度时间对齐能力,同时保持模态间粗粒度语义一致性。代码已公开于https://github.com/bbvisual/MoBind。

原文摘要 · Abstract (English)

We aim to learn a joint representation between inertial measurement unit (IMU) signals and 2D pose sequences extracted from video, enabling accurate cross-modal retrieval, temporal synchronization, subject and body-part localization, and action recognition. To this end, we introduce MoBind, a hierarchical contrastive learning framework designed to address three challenges: (1) filtering out irrelevant visual background, (2) modeling structured multi-sensor IMU configurations, and (3) achieving fine-grained, sub-second temporal alignment. To isolate motion-relevant cues, MoBind aligns IMU signals with skeletal motion sequences rather than raw pixels. We further decompose full-body motion into local body-part trajectories, pairing each with its corresponding IMU to enable semantically grounded multi-sensor alignment. To capture detailed temporal correspondence, MoBind employs a hierarchical contrastive strategy that first aligns token-level temporal segments, then fuses local (body-part) alignment with global (body-wide) motion aggregation. Evaluated on mRi, TotalCapture, and EgoHumans, MoBind consistently outperforms strong baselines across all four tasks, demonstrating robust fine-grained temporal alignment while preserving coarse semantic consistency across modalities. Code is available at https://github.com/bbvisual/ MoBind.

动作对齐多模态传感器融合姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。