arXiv:2510.02672eess.AScs.SD2025-10

用FiLM机制让语音变速模型更稳定,避免失真。

STSM-FiLM: A FiLM-Conditioned Neural Architecture for Time-Scale Modification of Speech

  • 用FiLM条件控制网络,根据速度因子动态调节语音特征。
  • 在多种编码器下均能保持语音清晰度,支持大幅变速。
  • 适合需要高质量语音变速的应用,如语音助手、字幕同步。

语音时长调整(TSM)旨在改变音频播放速度而不改变音高。传统方法如基于波形相似性的重叠相加(WSOLA)虽表现良好,但在非平稳或极端拉伸条件下常产生失真。本文提出STSM-FiLM——一种全神经架构,通过特征逐维线性调制(FiLM)对连续速度因子进行条件控制。模型通过监督学习,以WSOLA生成的输出为参考,学习其对齐与合成行为,同时利用深度学习提取的表征。我们测试了四种编码器-解码器组合:STFT-HiFiGAN、WavLM-HiFiGAN、Whisper-HiFiGAN和EnCodec,结果表明,STSM-FiLM可在广泛的时间缩放因子下生成感知一致的语音输出。整体表明,基于FiLM的条件控制能显著提升神经TSM模型的泛化能力与灵活性。

原文摘要 · Abstract (English)

Time-Scale Modification (TSM) of speech aims to alter the playback rate of audio without changing its pitch. While classical methods like Waveform Similarity-based Overlap-Add (WSOLA) provide strong baselines, they often introduce artifacts under non-stationary or extreme stretching conditions. We propose STSM-FILM - a fully neural architecture that incorporates Feature-Wise Linear Modulation (FiLM) to condition the model on a continuous speed factor. By supervising the network using WSOLA-generated outputs, STSM-FILM learns to mimic alignment and synthesis behaviors while benefiting from representations learned through deep learning. We explore four encoder-decoder variants: STFT-HiFiGAN, WavLM-HiFiGAN, Whisper-HiFiGAN, and EnCodec, and demonstrate that STSM-FILM is capable of producing perceptually consistent outputs across a wide range of time-scaling factors. Overall, our results demonstrate the potential of FiLM-based conditioning to improve the generalization and flexibility of neural TSM models.

语音处理神经网络时长调整FiLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。