构建首个合成语音假信息检测数据集,助力识别虚假语音内容
SpMis: An Investigation of Synthetic Spoken Misinformation Detection
- 构建包含1000+说话人、5类话题的合成语音数据集SpMis
- 利用先进语音合成技术生成高仿真语音,验证检测模型可行性
- 为防范语音诈骗提供基础研究支持,适合安全与可信AI研究者
近年来,语音生成技术在生成模型和大规模训练技术推动下迅速发展。尽管这提升了合成语音质量,但也带来了滥用风险,尤其是生成合成虚假信息。现有研究多聚焦于区分机器与人类语音,但更紧迫的问题是检测语音内容中的虚假信息。该任务需综合分析说话人身份、话题和合成方式等因素。为此,本文首次开展合成语音假信息检测研究,提出开源数据集SpMis,涵盖超过1000名说话人、五类常见话题,采用前沿文本转语音系统生成语音。实验显示检测具备潜力,但仍面临实际应用挑战,凸显该领域持续研究的重要性。
原文摘要 · Abstract (English)
In recent years, speech generation technology has advanced rapidly, fueled by generative models and large-scale training techniques. While these developments have enabled the production of high-quality synthetic speech, they have also raised concerns about the misuse of this technology, particularly for generating synthetic misinformation. Current research primarily focuses on distinguishing machine-generated speech from human-produced speech, but the more urgent challenge is detecting misinformation within spoken content. This task requires a thorough analysis of factors such as speaker identity, topic, and synthesis. To address this need, we conduct an initial investigation into synthetic spoken misinformation detection by introducing an open-source dataset, SpMis. SpMis includes speech synthesized from over 1,000 speakers across five common topics, utilizing state-of-the-art text-to-speech systems. Although our results show promising detection capabilities, they also reveal substantial challenges for practical implementation, underscoring the importance of ongoing research in this critical area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。