用0.09B参数模型实现跨模态生理信号高效生成,适合边缘设备部署。
Compact Latent Manifold Translation: A Parameter-Efficient Foundation Model for Cross-Modal and Cross-Frequency Physiological Signal Synthesis

- 采用两级离散翻译框架,用分层向量量化分离不同生理信号
- 跨模态生成时心电图峰值检测F1提升至0.83,频域超分辨率相关性达0.9956
- 参数极省,适合医疗边缘设备,可统一处理多模态、多频率信号
生理时间序列(如心电图ECG与光电容积脉搏波PPG)分析长期受异构设备带来的模态与频率差异阻碍。现有基础模型依赖连续潜在空间,常出现模态纠缠、高保真跨频生成能力不足,且计算开销大,难以部署于边缘设备。本文提出紧凑潜在流形翻译(CLMT),一个仅0.09B参数的统一框架,通过新颖的两阶段离散翻译范式解决上述问题。首先,引入通用分词器,利用分层残差向量量化(RVQ)将异构信号分解为独立、结构化的离散潜在流形,有效避免跨模态干扰;其次,上下文提示潜在翻译器通过融合静态生理先验,将离散标记在模态间映射,将复杂信号合成重构为纯潜在序列翻译任务。大量实验表明,0.09B模型显著优于大规模基线:在跨模态PPG转ECG合成中,解决时间相位漂移,临床R波检测F1分数从0.37提升至0.83;在极端跨频率超分辨率(25Hz至100Hz)下,成功恢复高频诊断特征,达到前所未有的皮尔逊相关性0.9956。通过以极小计算开销学习生物信号的通用离散语言,本方法为边缘可部署、多模态医疗基础模型开辟新路径。
原文摘要 · Abstract (English)
The analysis of physiological time series, such as electrocardiograms (ECG) and photoplethysmograms (PPG), is persistently hindered by modality and frequency gaps stemming from heterogeneous recording devices. Existing foundation models typically rely on continuous latent spaces, which frequently suffer from severe modality entanglement, lack high-fidelity cross-frequency generative capacity, and impose high computational costs that prohibit edge-device deployment. In this paper, we propose Compact Latent Manifold Translation (CLMT), a highly parameter-efficient (0.09B) unified framework that bridges these gaps through a novel two-stage discrete translation paradigm. First, we introduce a Universal Tokenizer utilizing Hierarchical Residual Vector Quantization (RVQ) to decouple heterogeneous signals into isolated, well-structured discrete latent manifolds, effectively preventing inter-modality interference. Second, a Context-Prompted Latent Translator maps these discrete tokens across modalities by integrating static physiological priors, reframing complex signal synthesis as a pure latent sequence translation task. Extensive evaluations demonstrate that our 0.09B model significantly outperforms massive baselines. In cross-modal PPG-to-ECG synthesis, it resolves temporal phase drift and dramatically improves the clinical R-peak detection F1-score from 0.37 (baseline) to 0.83. Furthermore, in extreme cross-frequency super-resolution (25Hz to 100Hz), it successfully recovers high-frequency diagnostic landmarks, achieving an unprecedented Pearson correlation of 0.9956. By learning a universal discrete language for biological signals with a fraction of the computational footprint, our approach sets a new trajectory for edge-deployable, multi-modal medical foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。