仅用一段普通语音即可生成任意人声的高保真ASMR,无需专门训练数据。
DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice
- 用大语言模型分离语义与ASMR风格,再通过流匹配重建音色。
- 单段普通语音可生成高质量ASMR,无需目标说话人的耳语数据。
- 适用于想快速生成个性化放松语音的开发者与内容创作者。
尽管现代文本转语音(TTS)系统在朗读式语音上已达到高保真度,但在生成自主感官高潮反应(ASMR)——一种低强度、专用于放松的特殊语音风格——方面仍面临挑战,主要因其细微特征常无发音,且需零样本说话人适配。本文提出首个针对零样本ASMR生成的框架DeepASMR。我们证明,仅需目标说话人一段普通的朗读语音片段,即可合成其本人的高保真ASMR语音,无需该说话人的专属耳语训练数据。方法上,我们发现离散语音标记能软性解耦ASMR风格与说话人音色。基于此,提出两阶段流程:利用大语言模型进行内容与风格编码,再通过流匹配声学解码器实现音色重建。此外,我们构建了包含670小时的中英文多说话人ASMR语音语料库DeepASMR-DB,引入融合客观指标、人工听评、大语言模型评分与无声语音分析的新评估协议。大量实验证明,DeepASMR在任意说话人生成的ASMR自然度与风格保真度上均达当前最优,同时保持正常语音合成的竞争力。
原文摘要 · Abstract (English)
While modern Text-to-Speech (TTS) systems achieve high fidelity for read-style speech, they struggle to generate Autonomous Sensory Meridian Response (ASMR), a specialized, low-intensity speech style essential for relaxation. The inherent challenges include ASMR's subtle, often unvoiced characteristics and the demand for zero-shot speaker adaptation. In this paper, we introduce DeepASMR, the first framework designed for zero-shot ASMR generation. We demonstrate that a single short snippet of a speaker's ordinary, read-style speech is sufficient to synthesize high-fidelity ASMR in their voice, eliminating the need for whispered training data from the target speaker. Methodologically, we first identify that discrete speech tokens provide a soft factorization of ASMR style from speaker timbre. Leveraging this insight, we propose a two-stage pipeline incorporating a Large Language Model (LLM) for content-style encoding and a flow-matching acoustic decoder for timbre reconstruction. Furthermore, we contribute DeepASMR-DB, a comprehensive 670-hour English-Chinese multi-speaker ASMR speech corpus, and introduce a novel evaluation protocol integrating objective metrics, human listening tests, LLM-based scoring and unvoiced speech analysis. Extensive experiments confirm that DeepASMR achieves state-of-the-art naturalness and style fidelity in ASMR generation for anyone of any voice, while maintaining competitive performance on normal speech synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。