arXiv:2506.03174cs.CVcs.AI2025-06被引 4

融合四模态数据,提升人体动作识别与跨模态检索精度

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks

  • 构建三视角视频+运动捕捉+IMU+文本的多模态融合模型
  • 零样本动作识别F1达0.6226,准确率73.2%,显著优于基线
  • 适合需要精细动作分析的医疗、机器人等场景

近年来,可穿戴设备的普及凸显了基于惯性测量单元(IMU)的行为分析重要性。尽管应用涵盖医疗、机器人等多个领域,现有研究日益转向多模态分析。已有工作提出结合第一人称视频与文本的多模态基础模型,但难以实现对全身动作的细致解析。为此,我们提出活动理解与表征对齐多模态基础模型(AURA-MFM),集成第三人称视频、运动捕捉数据、IMU信号和文本四类模态。通过引入第三人称视频与运动捕捉数据,模型能够实现比第一人称视角更全面的动作理解。同时,采用基于Transformer的IMU编码器进一步提升性能。在跨模态检索与动作识别任务上的实验表明,本模型优于现有方法。尤其在零样本动作分类中,本方法取得0.6226的F1分数和73.2%的准确率,而现有方法仅为0.0747和19.61%。

原文摘要 · Abstract (English)

In recent years, the widespread adoption of wearable devices has highlighted the growing importance of behavior analysis using IMU. While applications span diverse fields such as healthcare and robotics, recent studies have increasingly focused on multimodal analysis, in addition to unimodal analysis. Several studies have proposed multimodal foundation models that incorporate first-person video and text data; however, these models still fall short in providing a detailed analysis of full-body human activity. To address this limitation, we propose Activity Understanding and Representations Alignment - Multimodal Foundation Model (AURA-MFM), a foundational model integrating four modalities: third-person video, motion capture, IMU, and text. By incorporating third-person video and motion capture data, the model enables a detailed and multidimensional understanding of human activity, which first-person perspectives alone fail to capture. Additionally, a Transformer-based IMU encoder is employed to enhance the model's overall performance. Experimental evaluations on retrieval and activity recognition tasks demonstrate that our model surpasses existing methods. Notably, in the zero-shot classification for action recognition, our method achieved significantly higher performance, with an F1-score of 0.6226 and an accuracy of 0.7320, whereas the existing method recorded an F1-score of 0.0747 and an accuracy of 0.1961.

多模态动作识别IMU基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。