arXiv:2605.18837cs.LGcs.AI2026-05

提出VCR框架,让可穿戴设备在信号缺失时仍能稳定提取有效健康特征。

VCR: Learning Valid Contextual Representation for Incomplete Wearable Signals

论文配图:VCR: Learning Valid Contextual Representation for Incomplete Wearable Signals
图 1 · 摘自论文原文
  • 用正交令牌化分离共享语义与模态特有残差,避免信息冗余。
  • 仅重建缺失模态的共享部分,防止虚构不可推断的细节。
  • 在全模态、单缺失、多缺失场景下均优于现有方法。

可穿戴设备能从多模态信号实现持续健康监测,但真实部署受限于标注数据少和传感器普遍缺失。尽管大规模自监督预训练降低了对标签的依赖,多数现有方法假设所有模态均可用。当前处理模态缺失的方法常重建整个缺失信号,易产生无法从观测信号推断的模态特有细节,降低鲁棒性。本文提出VCR,一种自监督框架,旨在学习对模态缺失鲁棒的有效表示。VCR采用正交令牌化,通过修正潜在流形并应用几何投影,强制实现严格正交解耦,将每种模态分解为共享语义与模态特有残差。该设计在保持信息完整性的同时,为缺失情况下的稳健学习提供结构基础。所得令牌由感知缺失的专家混合模型处理,能适应不同模态可用模式。通过仅约束重建缺失模态的共享成分,VCR有效抑制了非可推断模态特有细节的幻觉。在多个健康监测任务中,相比强监督与自监督基线,VCR在全模态、单缺失及多缺失设置下均持续提升性能与鲁棒性。

原文摘要 · Abstract (English)

Wearable devices enable continuous health monitoring from multimodal signals, but real-world deployment is hindered by limited labeled data and pervasive sensor incompleteness. While large-scale self-supervised pretraining reduces label dependence, most existing methods assume full modality availability. Current approaches for handling modality missingness often reconstruct entire absent signals, which can encourage hallucinating modality-specific details that are not inferable from the observed sensor signals and degrade robustness. We propose VCR, a self-supervised framework that learns to extract valid representations robust to modality missingness. VCR employs an orthogonal tokenizer to enforce strict orthogonal disentanglement by rectifying latent manifolds and applying a geometric projection, separating each modality into shared semantics and modality-specific residuals. This design preserves complete information integrity while serving as a structural foundation for robust learning under modality missingness. The resulting tokens are processed by a missing-aware mixture-of-experts backbone that adapts to varying patterns of modality availability. By constraining the objective to reconstruct only the shared components of missing modalities, VCR effectively mitigates hallucinations of non-inferable modality-specific details. Across multiple health monitoring tasks, VCR consistently improves performance and robustness under full, single-missing, and multiple-missing modality settings compared with strong supervised and self-supervised baselines.

可穿戴健康自监督学习多模态缺失表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。