语音大模型能识别生理信号,多语言模型表现更优。
Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals
- 用语音模型提取生理信号时序特征,替代原始数据输入。
- 多语言语音模型在心电、肌电、皮电信号分类中准确率最高。
- 证明语音大模型可跨域应用于医疗信号分析,适合医学AI研究者。
尽管仅在语音数据上训练,语音基础模型(如Whisper)在音频分类等非语音任务中表现出色,这得益于语音与音频的共性。本研究将此类模型应用于更具挑战性的跨领域任务:分类生理时间序列信号。我们验证两个假设:其一,语音模型可通过捕捉共享的时间模式泛化到生理信号;其二,多语言语音模型因预训练中接触更多变异性,能生成更鲁棒的通用表征。实验基于心电图(ECG)、肌电图(EMG)和皮电反应(EDA)信号进行压力识别,结果表明,使用语音模型提取的特征优于直接使用原始生理信号的模型。其中,多语言语音模型表现最佳,支持了假设并展现了其在跨域任务中的潜力。本研究为语音基础模型在语音之外的新领域应用提供了实证支持。
原文摘要 · Abstract (English)
Despite being trained exclusively on speech data, speech foundation models (SFMs) like Whisper have shown impressive performance in non-speech tasks such as audio classification. This is partly because speech shares some common traits with audio, enabling SFMs to transfer effectively. In this study, we push the boundaries by evaluating SFMs on a more challenging out-of-domain (OOD) task: classifying physiological time-series signals. We test two key hypotheses: first, that SFMs can generalize to physiological signals by capturing shared temporal patterns; second, that multilingual SFMs will outperform others due to their exposure to greater variability during pre-training, leading to more robust, generalized representations. Our experiments, conducted for stress recognition using ECG (Electrocardiogram), EMG (Electromyography), and EDA (Electrodermal Activity) signals, reveal that models trained on SFM-derived representations outperform those trained on raw physiological signals. Among all models, multilingual SFMs achieve the highest accuracy, supporting our hypothesis and demonstrating their OOD capabilities. This work positions SFMs as promising tools for new uncharted domains beyond speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。