arXiv:2501.19034cs.CV2025-01中稿 · ACM IMWUT/UBICOMP …被引 18

用手机手表等设备的无线信号和传感器数据,实现室内动作自动识别与摘要。

XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses

  • 融合手机、手表等设备的Wi-Fi和惯性信号,构建多模态动作数据集。
  • 提出XRFMamba模型,在动作定位任务中达到78.74% mAP,参数减少35%。
  • 引入新评估指标RMC,动作摘要效果达0.802,适合智能健康与人机交互研究。

人体动作识别(HAR)在健康监测、智能家居和人机交互中至关重要。尽管HAR研究广泛,但利用智能家居环境中手机、手表、耳机和眼镜的Wi-Fi及惯性测量单元(IMU)信号进行连续动作识别与摘要仍属新兴任务。本文提出XRF V2数据集,用于室内日常活动的时间动作定位(TAL)与动作摘要。该数据集整合了来自16名志愿者在三种不同环境下的多模态数据,包括Wi-Fi信号、多种可穿戴设备的IMU数据及同步视频。为解决TAL与动作摘要问题,我们设计了XRFMamba神经网络,能有效捕捉未剪辑传感序列中的长期依赖,平均mAP达78.74,比最新WiFiTAD模型提升5.49个百分点,且参数量减少35%。在动作摘要任务中,我们引入新指标响应语义一致性(RMC),取得0.802的平均mRMC。XRF V2可推动动作定位、预测、姿态估计、多模态大模型预训练及合成数据生成等研究。数据与代码已开源:https://github.com/aiotgroup/XRFV2。

原文摘要 · Abstract (English)

Human Action Recognition (HAR) plays a crucial role in applications such as health monitoring, smart home automation, and human-computer interaction. While HAR has been extensively studied, action summarization using Wi-Fi and IMU signals in smart-home environments , which involves identifying and summarizing continuous actions, remains an emerging task. This paper introduces the novel XRF V2 dataset, designed for indoor daily activity Temporal Action Localization (TAL) and action summarization. XRF V2 integrates multimodal data from Wi-Fi signals, IMU sensors (smartphones, smartwatches, headphones, and smart glasses), and synchronized video recordings, offering a diverse collection of indoor activities from 16 volunteers across three distinct environments. To tackle TAL and action summarization, we propose the XRFMamba neural network, which excels at capturing long-term dependencies in untrimmed sensory sequences and achieves the best performance with an average mAP of 78.74, outperforming the recent WiFiTAD by 5.49 points in mAP@avg while using 35% fewer parameters. In action summarization, we introduce a new metric, Response Meaning Consistency (RMC), to evaluate action summarization performance. And it achieves an average Response Meaning Consistency (mRMC) of 0.802. We envision XRF V2 as a valuable resource for advancing research in human action localization, action forecasting, pose estimation, multimodal foundation models pre-training, synthetic data generation, and more. The data and code are available at https://github.com/aiotgroup/XRFV2.

动作识别多模态数据智能穿戴无线感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。