用多模态联邦学习提升动作识别准确率,保护隐私同时增强鲁棒性。
MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model
- 融合深度相机、压力垫和加速度计数据,用图注意力与专家混合模型统一表征。
- 在真实场景下达成中心化F1 0.934,联邦学习下仍保持0.881的高分。
- 适合关注隐私保护、多源数据协同的智能健康与智能家居开发者。
动作识别(HAR)在健身追踪、智慧家庭和医疗监测中至关重要。传统系统依赖单一模态(如运动传感器或摄像头),在真实环境中易受干扰。本文提出FedTime-MAGNET,一种结合深度相机、压力垫和加速度计的多模态联邦学习框架。核心为多模态自适应图神经专家变换器(MAGNET),利用图注意力与专家混合机制生成跨模态统一且具有区分性的嵌入表示。为捕捉复杂时序依赖,引入轻量级T5编码器架构。大量实验表明,该方法显著提升性能:中心化环境下F1得分为0.934,联邦学习下仍达0.881,验证了多模态融合、时序大模型与联邦学习结合的有效性,适用于构建精准且鲁棒的HAR系统。
原文摘要 · Abstract (English)
Human Activity Recognition (HAR) plays a vital role in applications such as fitness tracking, smart homes, and healthcare monitoring. Traditional HAR systems often rely on single modalities, such as motion sensors or cameras, limiting robustness and accuracy in real-world environments. This work presents FedTime-MAGNET, a novel multimodal federated learning framework that advances HAR by combining heterogeneous data sources: depth cameras, pressure mats, and accelerometers. At its core is the Multimodal Adaptive Graph Neural Expert Transformer (MAGNET), a fusion architecture that uses graph attention and a Mixture of Experts to generate unified, discriminative embeddings across modalities. To capture complex temporal dependencies, a lightweight T5 encoder only architecture is customized and adapted within this framework. Extensive experiments show that FedTime-MAGNET significantly improves HAR performance, achieving a centralized F1 Score of 0.934 and a strong federated F1 Score of 0.881. These results demonstrate the effectiveness of combining multimodal fusion, time series LLMs, and federated learning for building accurate and robust HAR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。