用语音文本联合表示学习,提升电声喉语音转换的自然度和可懂度。
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning

- 融合语音与文本信息,构建联合表示学习框架
- 在多个数据集上显著优于仅用语音的基线方法
- 适合语音康复、辅助沟通等临床应用场景
喉切除患者依赖电声喉设备发声。相比正常语音,电声喉语音存在严重失真、音素变化有限、语调不自然及时间偏移等问题,影响自然度和可懂度。尽管基于序列到序列(seq2seq)语音转换(VC)的电声喉语音转正常语音(EL2SP)方法有潜力,但电声喉语音与正常语音间的显著差异导致映射误差累积,限制性能。为此,本文提出一种新型表示学习框架,整合语音与文本表示,以提升seq2seq VC模型中的映射与重建质量。方法包括两阶段:1)表示融合与学习,使用预训练模块构建能融合辅助文本信息的网络,学习语音-文本联合表示;2)重建训练,采用自编码器式策略,在不增加模型复杂度的前提下完成最终建模。引入中层、输入层及混合层三种融合策略,逐步增强学习效果。此外,除标准的seq2seq VC目标外,还增加对联合表示的重建损失,以优化表示迁移。实验表明,在不同EL2SP数据集上,结合数据增强后,所提方法均显著优于仅依赖语音表示的基线。系统设计深度的逐步提升进一步验证了方法的有效性。本方法为电声喉语音增强及辅助通信技术提供了一种可扩展、实用的解决方案。
原文摘要 · Abstract (English)
Objective: laryngectomees depend on an electromechanical device to generate electrolaryngeal (EL) speech. Compared with normal speech, EL speech suffers from severe distortion, limited phonetic variation, unnatural prosody, and temporal shifts, degrading naturalness and intelligibility. Although sequence-to-sequence (seq2seq) voice conversion (VC) based EL-speech-to-normal-speech conversion (EL2SP) is promising, substantial mismatches between EL and normal speech inevitably cause cumulative mapping errors that limit performance. To address this, we describe a novel representation learning framework integrating speech and text representations to improve mapping and reconstruction quality within a seq2seq VC model. Methods: our methodology comprises two main stages: 1) representation integration and learning, and 2) reconstruction training. A network capable of incorporating auxiliary text information is first constructed with pretrained modules to learn speech--text-based integrated representations. Then, an autoencoder-style reconstruction strategy finalizes EL2SP model to inherit these representations without increasing model complexity. We introduce three fusion strategies including middle-, input-, and hybrid-level fusion strategies that progressively enhance learning. Moreover, besides standard seq2seq VC objectives, an additional reconstruction loss on the integrated representation is introduced to refine representation transfer. Results: experiments under different EL2SP datasets consistently demonstrate that our methods, combined with data augmentations, outperform baselines relying solely on speech representations. Furthermore, progressive improvements with system design depth validate the effectiveness of our methods. Significance: the proposed methods provide an extensible and practical methodology for EL speech enhancement and assistive communication technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。