arXiv:2504.02287cs.CV2025-04被引 4

构建家用多传感器动作识别数据集并提出动态融合新方法

MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion

  • 用分布式传感器采集带帧级标注的多模态视频
  • 在两个数据集上准确率超越现有最优方法
  • 适合做家庭场景智能监控与多视角融合研究者

多模态多视角动作识别是计算机视觉的快速发展的领域,具有显著的监控应用潜力。然而,现有数据集难以应对广域分布、异步数据流及缺乏帧级标注等现实挑战。同时,现有方法在建模跨视角关系和提升空间特征学习方面存在困难。本文提出MultiSensor-Home数据集,用于家庭环境下的全面动作识别,并引入基于Transformer的多模态多视角传感器融合方法(MultiTSF)。该数据集包含由分布式传感器捕获的未剪辑视频,提供高分辨率RGB和音频数据,以及详细的多视角帧级动作标签。MultiTSF方法采用Transformer融合机制,动态建模跨视角关系,并集成人体检测模块以增强空间特征学习,引导模型关注含人体活动的帧,从而提升识别精度。在所提MultiSensor-Home及已有MM-Office数据集上的实验表明,MultiTSF优于当前最优方法。定量与定性结果验证了该方法在真实世界多模态多视角动作识别中的有效性。源代码已公开于https://github.com/thanhhff/MultiTSF。

原文摘要 · Abstract (English)

Multi-modal multi-view action recognition is a rapidly growing field in computer vision, offering significant potential for applications in surveillance. However, current datasets often fail to address real-world challenges such as wide-area distributed settings, asynchronous data streams, and the lack of frame-level annotations. Furthermore, existing methods face difficulties in effectively modeling inter-view relationships and enhancing spatial feature learning. In this paper, we introduce the MultiSensor-Home dataset, a novel benchmark designed for comprehensive action recognition in home environments, and also propose the Multi-modal Multi-view Transformer-based Sensor Fusion (MultiTSF) method. The proposed MultiSensor-Home dataset features untrimmed videos captured by distributed sensors, providing high-resolution RGB and audio data along with detailed multi-view frame-level action labels. The proposed MultiTSF method leverages a Transformer-based fusion mechanism to dynamically model inter-view relationships. Furthermore, the proposed method integrates a human detection module to enhance spatial feature learning, guiding the model to prioritize frames with human activity to enhance action the recognition accuracy. Experiments on the proposed MultiSensor-Home and the existing MM-Office datasets demonstrate the superiority of MultiTSF over the state-of-the-art methods. Quantitative and qualitative results highlight the effectiveness of the proposed method in advancing real-world multi-modal multi-view action recognition. The source code is available at https://github.com/thanhhff/MultiTSF.

动作识别多模态融合家庭监控Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。