一种可统一处理多模态可穿戴数据的轻量级自编码框架,能解耦表征并生成真实数据。
Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations

- 通过多视图自监督分解损失与对称自编码器架构,学习跨模态时间-频率潜空间。
- 在30个模态的活动识别任务中,身份识别准确率提升6.75%,重建误差降低76.84%。
- 适合边缘智能穿戴设备与临床医疗场景,支持实时推理与多任务处理。
解耦表征学习是实现多模态可穿戴计算中通用、可持续模型的关键。然而,现有方法无法作为全栈可穿戴处理器,即未能同时解决特定任务分类性能、可解释的解耦表征学习、融合及高度异构多模态时序数据的生成建模问题。为填补这一空白,我们提出全景多模态变分分解自编码器(OmniDecVAEs),该框架能以统一且可扩展的方式从任意数量模态中高效学习多功能表征。OmniDecVAEs通过多视图自监督分解损失和共享非对称自编码器架构,学习模态条件的时间-频率潜子空间。在包含最多三十个模态的挑战性全景人体活动识别(HAR)设置中,结果显示OmniDecVAEs能够学习全栈可穿戴表征。相较于基于Transformer和VAE的方法,其解耦表征特性分别带来活动识别1.01%和身份识别6.75%的准确率提升。此外,OmniDecVAEs生成的全景时频数据表现出更优的重构效果(平均绝对误差降低76.84%)与真实/合成数据分布相似性(最大均值差异改善13.85%)。结果表明,OmniDecVAEs具有轻量化、高表示能力、模态无关的空间复杂度(410万参数)和实时延迟,适用于智能边缘可穿戴设备与临床健康应用,单一模型统一实现多种处理需求。
原文摘要 · Abstract (English)
Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal wearable computing. However, existing approaches do not operate as full-stack wearable processors, i.e., they do not simultaneously address task-specific classification performance, disentangled and interpretable representation learning, fusion, and generative modeling of highly heterogeneous multi-modal time series. To address this gap, we introduce Omni-modal Variational Decomposition Autoencoders (OmniDecVAEs), a framework that efficiently learns multi-purpose representations in a unified and scalable manner from arbitrarily many modalities. OmniDecVAEs extend DecVAEs by learning modality-conditioned time-frequency latent subspaces through a multi-view self-supervised decomposition loss and a shared asymmetric autoencoder (AE) architecture. Results on a challenging omni-modal human activity recognition (HAR) setting with up to thirty modalities, demonstrate the ability of OmniDecVAEs to learn full-stack wearable representations. When compared to transformer-based and VAE-based methods, OmniDecVAEs full-stack disentangled representation properties lead to accuracy improvements of 1.01% and 6.75% in activity and identity recognition, respectively. Furthermore, OmniDecVAEs synthesize realistic omni-modal time-frequency data that manifest with enhanced reconstructions (mean absolute error improves by 76.84%) and distributional similarity between real and synthetic data (maximum mean discrepancy improves by 13.85%). Our results highlight OmniDecVAEs potential as a lightweight model suitable for intelligent edge wearables and clinical healthcare, unifying processing requirements and abilities in a single model, through its enhanced representational capacity, modality-invariant spatial complexity (4.1M parameters), and real-time latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。