arXiv:2504.02279cs.CV2025-04被引 3

用Transformer融合多视角多模态数据,提升人体动作识别准确率

MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition

  • 基于Transformer动态建模多视角关系与时间依赖
  • 在两个数据集上均超越当前最优方法,视频级与帧级识别更准
  • 自动生成伪标签,减少人工标注,适合智能监控场景

从多模态和多视角观测中进行动作识别在安防、机器人和智慧环境中有重要应用前景。然而现有方法难以应对真实场景中的多样环境条件、严格传感器同步要求以及细粒度标注需求。本文提出多模态多视角Transformer传感器融合方法(MultiTSF),利用Transformer动态建模跨视角关系并捕捉多视角间的时序依赖。此外,引入人体检测模块生成伪真值标签,使模型聚焦含人体活动的帧,增强空间特征学习。在自建MultiSensor-Home数据集和现有MM-Office数据集上的全面实验表明,MultiTSF在视频序列级和帧级动作识别任务中均优于当前最优方法。

原文摘要 · Abstract (English)

Action recognition from multi-modal and multi-view observations holds significant potential for applications in surveillance, robotics, and smart environments. However, existing methods often fall short of addressing real-world challenges such as diverse environmental conditions, strict sensor synchronization, and the need for fine-grained annotations. In this study, we propose the Multi-modal Multi-view Transformer-based Sensor Fusion (MultiTSF). The proposed method leverages a Transformer-based to dynamically model inter-view relationships and capture temporal dependencies across multiple views. Additionally, we introduce a Human Detection Module to generate pseudo-ground-truth labels, enabling the model to prioritize frames containing human activity and enhance spatial feature learning. Comprehensive experiments conducted on our in-house MultiSensor-Home dataset and the existing MM-Office dataset demonstrate that MultiTSF outperforms state-of-the-art methods in both video sequence-level and frame-level action recognition settings.

动作识别多模态融合Transformer传感器融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。