arXiv:2601.08204cs.CV2026-01

用可穿戴设备和无线信号自动生成日常活动的自然语言描述

MobiDiary: Autoregressive Action Captioning with Wearable Devices and Wireless Signals

  • 统一编码器融合惯性与无线信号,捕捉运动动态特征
  • 在多个公开数据集上达到当前最优的描述生成效果
  • 适合健康监测与智能居家场景,无需摄像头保障隐私

智能家庭中的人体活动识别对健康监测和辅助生活至关重要。尽管视觉系统普遍使用,但面临隐私担忧和环境限制(如遮挡)。本文提出 MobiDiary 框架,直接从异构物理信号(特别是 IMU 和 Wi-Fi)生成自然语言活动描述。不同于传统方法仅输出预定义标签,MobiDiary 生成表达丰富、可读性强的摘要。为弥合连续噪声信号与离散语言描述之间的语义鸿沟,我们设计统一传感器编码器,利用片块机制捕捉局部时序相关性,并通过异构放置嵌入统一不同传感器的空间上下文。这些统一信号标记输入基于 Transformer 的自回归解码器,逐词生成连贯动作描述。我们在多个公开基准(XRF V2、UWash、WiFiTAD)上全面评估,结果表明 MobiDiary 能有效跨模态泛化,在描述生成指标(如 BLEU@4、CIDEr、RMC)上表现领先,优于专用基线模型,在连续活动理解任务中展现卓越性能。

原文摘要 · Abstract (English)

Human Activity Recognition (HAR) in smart homes is critical for health monitoring and assistive living. While vision-based systems are common, they face privacy concerns and environmental limitations (e.g., occlusion). In this work, we present MobiDiary, a framework that generates natural language descriptions of daily activities directly from heterogeneous physical signals (specifically IMU and Wi-Fi). Unlike conventional approaches that restrict outputs to pre-defined labels, MobiDiary produces expressive, human-readable summaries. To bridge the semantic gap between continuous, noisy physical signals and discrete linguistic descriptions, we propose a unified sensor encoder. Instead of relying on modality-specific engineering, we exploit the shared inductive biases of motion-induced signals--where both inertial and wireless data reflect underlying kinematic dynamics. Specifically, our encoder utilizes a patch-based mechanism to capture local temporal correlations and integrates heterogeneous placement embedding to unify spatial contexts across different sensors. These unified signal tokens are then fed into a Transformer-based decoder, which employs an autoregressive mechanism to generate coherent action descriptions word-by-word. We comprehensively evaluate our approach on multiple public benchmarks (XRF V2, UWash, and WiFiTAD). Experimental results demonstrate that MobiDiary effectively generalizes across modalities, achieving state-of-the-art performance on captioning metrics (e.g., BLEU@4, CIDEr, RMC) and outperforming specialized baselines in continuous action understanding.

活动识别自然语言生成可穿戴设备多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。