arXiv:2506.11620cs.SDcs.CL2025-06被引 1

用AI模拟听力损失,自动设计更精准的语音测试

(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test

  • 用ASR系统模拟听觉损失,识别常见发音混淆模式
  • 筛选出25对最优语音词对,诊断效能超越传统指标
  • 适合听力损伤研究者快速开发高精度测试工具

传统测听无法充分刻画听力损失对语音理解的功能影响,尤其针对老年性听力下降常见的阈上缺陷。为此,我们提出模拟音素语音测试(SimPhon Speech Test)方法,一种多阶段计算流程,用于在硅环境中设计与验证音素平衡的最小对语音测试。该方法利用现代自动语音识别(ASR)系统作为人类听觉的代理,通过可控声学退化处理语音刺激,首先识别最常见的音素混淆模式。这些模式指导从大型语言语料库中数据驱动地筛选候选词对。后续经过模拟诊断测试、专家人工筛选及最终针对性敏感性分析,逐步精简为最终优化的25对(SimPhon Speech Test-25)。关键发现是,该测试项目诊断性能与标准语音可懂度指数(SII)无显著相关性,表明其捕捉了超出简单可听性的感知缺陷。此计算优化测试集显著提升听力测试开发效率,已具备开展初步人体试验条件。

原文摘要 · Abstract (English)

Traditional audiometry often provides an incomplete characterization of the functional impact of hearing loss on speech understanding, particularly for supra-threshold deficits common in presbycusis. This motivates the development of more diagnostically specific speech perception tests. We introduce the Simulated Phoneme Speech Test (SimPhon Speech Test) methodology, a novel, multi-stage computational pipeline for the in silico design and validation of a phonetically balanced minimal-pair speech test. This methodology leverages a modern Automatic Speech Recognition (ASR) system as a proxy for a human listener to simulate the perceptual effects of sensorineural hearing loss. By processing speech stimuli under controlled acoustic degradation, we first identify the most common phoneme confusion patterns. These patterns then guide the data-driven curation of a large set of candidate word pairs derived from a comprehensive linguistic corpus. Subsequent phases involving simulated diagnostic testing, expert human curation, and a final, targeted sensitivity analysis systematically reduce the candidates to a final, optimized set of 25 pairs (the SimPhon Speech Test-25). A key finding is that the diagnostic performance of the SimPhon Speech Test-25 test items shows no significant correlation with predictions from the standard Speech Intelligibility Index (SII), suggesting the SimPhon Speech Test captures perceptual deficits beyond simple audibility. This computationally optimized test set offers a significant increase in efficiency for audiological test development, ready for initial human trials.

语音测试听力评估AI模拟ASR应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。