arXiv:2607.04154cs.SDcs.AI2026-07

用几何方法区分AI合成语音与真人发音,基于日语元音分布差异。

Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese)

  • 将语音谱视为耳蜗接收的频带概率分布,用Wasserstein距离度量
  • AI合成语音元音间距离短,自然语音分布更分散
  • 结合持久同调拓扑分析,可有效聚类区分两类语音

本文提出一种新方法,用于区分由生成式AI合成的假人声与真实语音。由于合成语音基于有限训练频谱生成,其元音多样性受限;而真实语音因人类发音器官灵活,元音频谱分布更丰富。以日语为例(五种元音一一对应特定音素),通过将语音谱归一化为耳蜗毛细胞接收频带的概率密度函数,利用Wasserstein距离衡量元音间差异。结果显示,合成语音的元音间Wasserstein距离较短。通过保持该距离并使用持久同调进行拓扑映射,可将合成与自然语音的谱概率密度函数分解为不同簇,实现有效区分。

原文摘要 · Abstract (English)

This paper explains the principles and provides examples of a new method for distinguishing between FAKE human speech synthesized by generative AI and natural speech. Since synthetic speech is generated based on information from a limited set of training spectra, the variety of vowels - which are key to identifying individuals - is limited. In contrast, natural speech exhibits a more diverse distribution of vowel spectra due to the flexibility of the human articulatory organ. In this paper, using Japanese - a Syllabary limited to five vowel phonemes, each of which corresponds one-to-one with a specific sound - as an example, we outline a method for distinguishing between synthetic and natural speech reading the same text by analyzing the spectral distributions. If we normalize the spectra of speech sounds and regard them as probability density functions for the frequency bands received by the hair cells of the human cochlea, and evaluate the distance between spectra using the Wasserstein metric, the Wasserstein distances between the vowels of synthetic speech are short. By preserving this distance and performing a topological mapping using persistent homology, the spectral probability density functions of synthetic and natural speech can be decomposed into clusters.

语音识别生成模型信息几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。