arXiv:2606.27886cs.LG2026-06

对比七种融合方法在多模态动作识别中的表现,发现门控融合效果最佳。

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset

论文配图:A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset
图 1 · 摘自论文原文
  • 采用七种主流融合策略比较性能,统一测试于HARMES数据集。
  • 门控融合模型达到0.82宏F1,比基线高6个百分点。
  • 适合关注多模态传感器融合的智能穿戴与健康监测研究者。

可穿戴传感器在人体活动识别(HAR)中的最新进展表明,多模态深度学习模型始终优于单模态模型。模态可包括加速度计(IMU)、RGB相机、音频信号等。多模态深度学习的一个关键环节是传感器融合方式的选择。近年来,已有多种融合范式被提出用于多模态HAR,但据我们所知,尚无在统一基准数据集上对这些范式的直接比较。为填补这一空白,本文在新发布的HARMES数据集上系统性地比较了七种前沿的传感器融合方法。该数据集包含61小时的完整标注数据,涵盖IMU、音频和环境湿度三种模态,聚焦15类日常生活活动(ADL)。通过将七种不同融合技术应用于一个先进的多模态模型架构,我们发现门控多模态融合(Gated Multi-modal Fusion)在留一人外交叉验证下取得最高宏F1分数0.82,相比基于拼接的晚融合基线(0.76)提升6个百分点。所有实验代码已公开在GitHub。

原文摘要 · Abstract (English)

Recent advances in Human Activity Recognition (HAR) from wearable sensors have shown that multi-modal deep learning models consistently outperform their uni-modal counterparts. Modalities can include IMUs, RGB cameras, audio signals, and others. One important aspect of multi-modal deep learning is the sensor fusion approach we apply. Over recent years, multiple fusion paradigms have been proposed for multi-modal HAR. However, to the best of our knowledge, no head-to-head comparison of these paradigms exists on a common multi-modal HAR benchmark dataset. To address this research gap, we systematically compare seven state-of-the-art sensor fusion methods on the recently released HARMES dataset, which comprises 61 hours of fully labeled IMU, audio, and ambient humidity data. The chosen dataset focuses on 15 household and personal hygiene activities of daily living (ADLs). By applying the seven different fusion techniques to a state-of-the-art multi-modal model architecture, we show that Gated Multi-modal Fusion achieves the highest macro F1-score (0.82), surpassing the concatenation-based late fusion HARMES paper baseline of 0.76 by +6pp under leave-one-participant-out evaluation. All code used in our experiments is made publicly available on GitHub.

动作识别多模态融合可穿戴设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。