用统一模型实现语音分析、控制与生成,支持精准调节音高、说话人等属性。
AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder
- 基于掩码自编码器,一模型完成语音分析与生成
- 可准确估计音高、说话人、信噪比等6项关键属性
- 适合语音合成、语音增强与个性化语音控制场景
本文提出AnCoGen,一种基于掩码自编码器的新型方法,将语音分析、控制与生成统一于单一模型中。AnCoGen可分析语音并估计说话人身份、音高、内容、音量、信噪比及清晰度指数等关键属性;同时可根据这些属性生成语音,并通过修改属性实现对合成语音的精确控制。大量实验验证了AnCoGen在语音分析-重合成、音高估计、音高调整及语音增强任务中的有效性。
原文摘要 · Abstract (English)
This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key attributes, such as speaker identity, pitch, content, loudness, signal-to-noise ratio, and clarity index. In addition, it can generate speech from these attributes and allow precise control of the synthesized speech by modifying them. Extensive experiments demonstrated the effectiveness of AnCoGen across speech analysis-resynthesis, pitch estimation, pitch modification, and speech enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。