arXiv:2505.14206cs.LGcs.AI2025-05

评估生成模型在可穿戴传感器数据合成中的局限性,揭示其在多模态与长时序上的不足。

Challenges and Limitations of Generative AI in Synthesizing Wearable Sensor Data

  • 系统评测生成对抗网络与扩散模型在复杂场景下的表现
  • 发现现有模型难以保持跨模态一致性与时间连贯性
  • 适合关注可穿戴设备数据生成的科研人员与工程师

可穿戴传感器的广泛应用有望提供海量异构的时间序列数据,推动人工智能在人体感知中的应用。然而,由于严格的伦理规范、隐私顾虑及其他限制,数据采集仍受限,阻碍了该领域的发展。通过生成对抗网络和扩散模型进行合成数据生成,被视为缓解数据稀缺与隐私问题的潜在解决方案。但这些模型通常仅适用于短时、单模态信号等狭窄场景。为填补这一空白,本文对当前最先进的时间序列生成模型进行了系统性评估,重点考察其在压力与情绪识别等挑战性任务中的表现。研究分析了模型在处理多模态、捕捉长程依赖以及支持条件生成方面的能力——这三者是真实可穿戴数据生成的核心需求。为实现公平严谨的比较,我们引入了一个评估框架,同时衡量生成数据的内在保真度及其在下游预测任务中的实用性。结果表明,现有方法在跨模态一致性、时间连贯性维持以及在‘用合成数据训练、真实数据测试’和数据增强场景下的鲁棒性方面存在显著缺陷。最后,我们提出了未来研究方向,以提升时间序列合成质量,并增强生成模型在可穿戴计算领域的适用性。

原文摘要 · Abstract (English)

The widespread adoption of wearable sensors has the potential to provide massive and heterogeneous time series data, driving the use of Artificial Intelligence in human sensing applications. However, data collection remains limited due to stringent ethical regulations, privacy concerns, and other constraints, hindering progress in the field. Synthetic data generation, particularly through Generative Adversarial Networks and Diffusion Models, has emerged as a promising solution to mitigate both data scarcity and privacy issues. However, these models are often limited to narrow operational scenarios, such as short-term and unimodal signal patterns. To address this gap, we present a systematic evaluation of state-of-the-art generative models for time series data, explicitly assessing their performance in challenging scenarios such as stress and emotion recognition. Our study examines the extent to which these models can jointly handle multi-modality, capture long-range dependencies, and support conditional generation-core requirements for real-world wearable sensor data generation. To enable a fair and rigorous comparison, we also introduce an evaluation framework that evaluates both the intrinsic fidelity of the generated data and their utility in downstream predictive tasks. Our findings reveal critical limitations in the existing approaches, particularly in maintaining cross-modal consistency, preserving temporal coherence, and ensuring robust performance in train-on-synthetic, test-on-real, and data augmentation scenarios. Finally, we present our future research directions to enhance synthetic time series generation and improve the applicability of generative models in the wearable computing domain.

生成模型可穿戴数据时间序列多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。