arXiv:2505.23378cs.LG2025-05中稿 · Interspeech 2025

用元学习建模语音疲劳,提升健康监测精度。

Meta-Learning Approaches for Speaker-Dependent Voice Fatigue Models

  • 将语音疲劳建模转化为元学习任务,避免重复训练。
  • 基于10,286段录音,变压器模型预测准确率最高。
  • 适合语音健康监测与个性化医疗研究者。

说话人依赖建模可显著提升语音健康监测性能。尽管混合效应模型常用于说话人适应,但每次新观测都需耗时重训练,难以投入生产。本文将该任务重构为元学习问题,探索三种逐步复杂的方案:基于集成的距离模型、原型网络和基于变压器的序列模型。利用预训练语音嵌入,在包含1,185名轮班工人、10,286段录音的大规模纵向数据集上,以语音中睡眠后时间作为疲劳指标进行预测。结果表明,所有测试的元学习方法均优于横断面及传统混合效应模型,其中基于变压器的方法表现最佳。

原文摘要 · Abstract (English)

Speaker-dependent modelling can substantially improve performance in speech-based health monitoring applications. While mixed-effect models are commonly used for such speaker adaptation, they require computationally expensive retraining for each new observation, making them impractical in a production environment. We reformulate this task as a meta-learning problem and explore three approaches of increasing complexity: ensemble-based distance models, prototypical networks, and transformer-based sequence models. Using pre-trained speech embeddings, we evaluate these methods on a large longitudinal dataset of shift workers (N=1,185, 10,286 recordings), predicting time since sleep from speech as a function of fatigue, a symptom commonly associated with ill-health. Our results demonstrate that all meta-learning approaches tested outperformed both cross-sectional and conventional mixed-effects models, with a transformer-based method achieving the strongest performance.

语音健康元学习疲劳检测个性化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。