arXiv:2412.18839cs.SDcs.AI2024-12中稿 · IEEE ICASSP 2025被引 4

用新方法和多模态数据提升无声低语转语音的清晰度与泛化能力

Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset

  • 通过音素对齐学习,直接从无声低语生成语音,减少对真实低语依赖
  • 引入基于扩散模型的唇动转语音技术,提升语音生成质量与跨说话人泛化性
  • 发布含7.96小时多模态数据的MultiNAM数据集,支持方法对比与评估

当前无声低语(NAM)转语音技术依赖语音克隆,从配对低语中模拟真实语音,但生成语音常不清晰且跨说话人泛化差。本文聚焦于从配对低语与文本中学习音素级对齐,并使用文语转换(TTS)系统生成语音。为降低对低语或口语音频的依赖,我们直接从NAM中学习音素对齐,但受限于训练数据质量。为进一步减少对原始数据的依赖,提出结合唇部动作信息推断语音,并引入基于扩散模型的新方法,利用近期唇动转语音技术进展。此外,我们发布了包含超过7.96小时配对的NAM、低语、视频与文本数据的MultiNAM数据集,并在该数据集上对所有方法进行基准测试。语音样本与数据集可在https://diff-nam.github.io/DiffNAM/获取。

原文摘要 · Abstract (English)

Current Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intelligibility and fails to generalize well across different speakers. To address this issue, we focus on learning phoneme-level alignments from paired whispers and text and employ a Text-to-Speech (TTS) system to simulate the ground-truth. To reduce dependence on whispers, we learn phoneme alignments directly from NAMs, though the quality is constrained by the available training data. To further mitigate reliance on NAM/whisper data for ground-truth simulation, we propose incorporating the lip modality to infer speech and introduce a novel diffusion-based method that leverages recent advancements in lip-to-speech technology. Additionally, we release the MultiNAM dataset with over 7.96 hours of paired NAM, whisper, video, and text data from two speakers and benchmark all methods on this dataset. Speech samples and the dataset are available at https://diff-nam.github.io/DiffNAM/

语音生成多模态扩散模型无声语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。