arXiv:2510.21317eess.AScs.SD2025-10被引 3

用语言模型检测语音生成中的胡言乱语,提升无参考评估的可靠性。

Are These Even Words? Quantifying the Gibberishness of Generative Speech Models

  • 基于语言模型在无参考条件下识别合成语音中的荒谬语句。
  • 发布高质量胡言乱语语音数据集,支持评估方法研发。
  • 适合语音质量评估与生成模型可信度研究者使用。

当前大量研究聚焦于非侵入式语音质量与可懂度评估,尤其在构建大规模真实场景语音数据集时至关重要。然而,随着生成模型合成高质量语音能力的提升,新型伪影如生成幻觉日益突出。尽管侵入式指标可从参考信号中发现此类差异,但现有非侵入方法对高质量音素混淆或极端胡言乱语的响应尚不明确。本文在完全无监督设置下,利用语言模型探索这一问题。同时,我们公开了一个高质量合成胡言乱语语音数据集,并提供多种语音语言模型计算评分的代码,以推动对口语中不合理句子评估方法的发展。

原文摘要 · Abstract (English)

Significant research efforts are currently being dedicated to non-intrusive quality and intelligibility assessment, especially given how it enables curation of large scale datasets of in-the-wild speech data. However, with the increasing capabilities of generative models to synthesize high quality speech, new types of artifacts become relevant, such as generative hallucinations. While intrusive metrics are able to spot such sort of discrepancies from a reference signal, it is not clear how current non-intrusive methods react to high-quality phoneme confusions or, more extremely, gibberish speech. In this paper we explore how to factor in this aspect under a fully unsupervised setting by leveraging language models. Additionally, we publish a dataset of high-quality synthesized gibberish speech for further development of measures to assess implausible sentences in spoken language, alongside code for calculating scores from a variety of speech language models.

语音生成语言模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。