只需一个样本,即可跨语言分离目标人声,无需训练。
Wanna hear your voice? A sample is all we need!
- 用频调门控机制动态调整目标语音特征,减少对语言特性的依赖。
- 零样本跨语言效果领先:越南语数据达14.8 dB,英语测试超13.8 dB。
- 适合低资源语言语音分离,无需标注数据或微调。
基于音频线索的目标说话人分离(TSE)研究长期聚焦于混合信号与参考语音建模,在英语数据上已取得优异成果。然而跨语言特性仍待深入,低资源语言受限于标注数据和语言资源匮乏。为此,我们提出WHYV(Wanna Hear Your Voice)框架,实现无需微调的零样本跨语言适应。WHYV采用频率调制门控机制,动态调节目标说话人声学特征,降低对语言特定线索的依赖。评估显示其达到当前最优零样本性能:在Libri2Mix mix-both上为13.8 dB,mix-clean上为18.1 dB,越南语数据上达14.8 dB。
原文摘要 · Abstract (English)
Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. However, cross-lingual properties remain underexplored, as low-resource languages face challenges from limited annotated data and linguistic resources. To bridge this gap, we propose WHYV (Wanna Hear Your Voice), a cross-lingual TSE framework enabling zero-shot adaptation without fine-tuning. WHYV employs a frequency-modulated gating mechanism that dynamically adjusts the acoustic features of the target speaker, minimizing reliance on language-specific cues. Evaluations demonstrate state-of-the-art zero-shot performance: 13.8 dB (Libri2Mix mix-both), 18.1 dB (mix-clean), and 14.8 dB on Vietnamese data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。