arXiv:2510.13068cs.LGcs.AI2025-10被引 5

用多尺度分解与相位感知训练,提升脑电心电肌电信号的生成质量。

NeuroRVQ: Multi-Scale Biosignal Tokenization for Generative Foundation Models

  • 通过多尺度卷积分解信号,分频段编码并保留高频细节。
  • 在脑电、心电、肌电上实现高保真重建,优于现有模型。
  • 适合需要精确生理信号建模的研究者,如神经科学与医疗AI。

脑电信号(EEG)、心电信号(ECG)和肌电信号(EMG)跨越多个时间与频域尺度,蕴含丰富的生理信息,但对机器学习而言难以处理。基于掩码信号令牌预测的基座模型虽展现出学习通用表示的潜力,但其性能依赖于分词器对高频动态的保持能力及高保真重建效果。本文提出NeuroRVQ,一种针对多模态生物信号的自适应分词器家族,通过多尺度时序卷积将信号分解为频段特异性表征,每类表征使用分层矢量量化(RVQ)代码本编码以保留高频细节,并引入新颖的相位感知训练损失,尊重傅里叶相位的环形拓扑结构。通过调节时序分辨率、卷积核数量与大小、以及RVQ深度,该设计可适配不同生物信号的谱-时特性。为验证分词器质量对下游任务的影响,我们为各模态训练了简单的掩码令牌基座模型(NeuroRVQ-FM)。结果表明,NeuroRVQ-FM在脑电、心电、肌电任务中表现优于或媲美现有专用基座模型,证明高保真分词是有效生物信号建模的关键因素。

原文摘要 · Abstract (English)

Biosignals such as electroencephalography (EEG), electrocardiography (ECG), and electromyography (EMG) encode physiological activity across multiple temporal and spectral scales, yielding representations that are rich but challenging for machine learning. Foundation models trained to predict masked signal tokens have shown promise in learning generalizable biosignal representations, yet their performance depends on the tokenizer's ability to preserve high-frequency dynamics and reconstruct signals with high fidelity. We introduce NeuroRVQ, a modality-adaptive biosignal tokenizer family designed for high-fidelity signal reconstruction. To capture the full frequency spectrum, NeuroRVQ decomposes biosignals into frequency-specific representations via multi-scale temporal convolutions, each encoded into hierarchical RVQ codebooks to preserve high-frequency detail, combined with a novel phase-aware training loss that respects the circular topology of Fourier phase. By tuning the temporal resolution, number and size of temporal kernels and RVQ depth, this design adapts to the spectro-temporal characteristics of each biosignal modality. To validate that tokenizer quality drives downstream performance, we train a simple masked-token foundation model for each modality (NeuroRVQ-FM) using the corresponding NeuroRVQ tokenizer. The NeuroRVQ-FM family achieves competitive or superior downstream performance compared to existing modality-specific foundation models, demonstrating that high-fidelity tokenization is a critical factor for effective biosignal modeling.

生物信号生成模型分词器神经信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。