arXiv:2608.24707cs.CL2026-08

首个跨语言语音幻觉检测基准,覆盖英俄哈三语真实与合成数据。

Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

论文配图:Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts
图 1 · 摘自论文原文
  • 构建12013条三语新闻样本,含可控幻觉类型与严重度的文本音频对。
  • 基于原文的检测效果优于直接音频处理,合成训练模型在真实假新闻上F1达0.82-0.88。
  • 揭示合成数据中模型风格信号干扰真实性判断,警示基准设计陷阱。

尽管文本幻觉检测已有广泛研究,但语音幻觉检测仍基本空白,尤其在低资源语言中。本文提出首个多语言语音幻觉检测基准,包含英、俄、哈三语共12,013条新闻样本,涵盖三种幻觉类型与三个严重等级。样本包括原始文章及对应的文本与音频幻觉版本。此外,补充了290条俄语(225条)和哈萨克语(65条)原生采集的已核查假新闻,经翻译并经相同TTS-ASR流程生成。评估微调的多语言编码器及零样本上下文的多模态解码器在基于转录本与直接音频处理上的表现。转录本检测普遍优于直接音频处理,强编码器在二分类任务中的性能下降与各语言的ASR错误率相关。在真实假新闻上,合成训练的检测器迁移表现良好(原文字典上宏F1 0.82–0.88),俄语来源分析显示存在与真实性相关及模型依赖的机器风格信号,量化了合成幻觉基准中的关键混淆因素。

原文摘要 · Abstract (English)

While text-based hallucination detection has been extensively studied, spoken hallucination detection remains largely unexplored, particularly for low-resource languages. We present the first multilingual spoken hallucination benchmark comprising 12,013 news samples across English, Russian, and Kazakh with controlled hallucinations of three types and three severity levels. Samples comprise original articles and aligned hallucinated counterparts in text and audio. We complement the synthetic corpus with 290 fact-checked fake news items collected natively in Russian (225) and Kazakh (65), translated into the other language and rendered through the same TTS-ASR pipeline. We assess fine-tuned multilingual encoders and, in zero-shot in-context settings, multimodal decoder models on transcript-based versus direct audio processing. Transcript-based detection generally outperforms direct audio processing, with binary-task degradation for strong encoders tracking per-language ASR error. On real-world fakes, synthetic-trained detectors transfer strongly (macro-F1 0.82-0.88 on original text), while Russian provenance analysis reveals both veracity-related and model-dependent machine-style signals, quantifying a key confound in synthetic hallucination benchmarks.

语音幻觉多语言虚假信息基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。