构建名人语音伪造数据集,提升合成语音真实度与检测挑战性。
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges
- 自动化采集+基于转录的分割,提升真声样本质量
- 合成语音自然度达NISQA-TTS 3.69分,人眼误判率61.9%
- 适用于语音伪造检测与防御研究,适合安全与可信AI领域
语音合成技术的进展带来了维护声音真实性的严峻挑战,尤其针对常被冒用的公众人物。本文提出一套完整的收集、筛选与生成合成语音数据的方法,聚焦政治人物,并详述实施中的难点。我们设计了自动化流程以获取高质量真声样本,采用基于转录的分割方法显著提升了合成语音质量。实验对比了单说话人到零样本合成等多种技术路径,记录了方法演进过程。最终数据集包含十位公众人物的真声与合成语音样本,合成语音自然度评分(NISQA-TTS)达3.69,人类误判率最高为61.9%。
原文摘要 · Abstract (English)
Recent advances in speech synthesis have introduced unprecedented challenges in maintaining voice authenticity, particularly concerning public figures who are frequent targets of impersonation attacks. This paper presents a comprehensive methodology for collecting, curating, and generating synthetic speech data for political figures and a detailed analysis of challenges encountered. We introduce a systematic approach incorporating an automated pipeline for collecting high-quality bonafide speech samples, featuring transcription-based segmentation that significantly improves synthetic speech quality. We experimented with various synthesis approaches; from single-speaker to zero-shot synthesis, and documented the evolution of our methodology. The resulting dataset comprises bonafide and synthetic speech samples from ten public figures, demonstrating superior quality with a NISQA-TTS naturalness score of 3.69 and the highest human misclassification rate of 61.9\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。