arXiv:2509.03021eess.AScs.SD2025-09中稿 · IEEE ICCE-TW 2025

用大模型实现无需训练的助听器语音可懂度评估

A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models

  • 基于大模型构建个性化助听器语音评估框架
  • 相比原模型误差降低2.59%,提升评估精度
  • 适合听力辅助设备研发与个性化调试

本研究聚焦于利用大语言模型(LLMs)实现助听器(HA)的零样本非侵入式语音可懂度评估。提出GPT-Whisper-HA,作为GPT-Whisper的扩展,该模型结合MSBG听力损失模型和NAL-R仿真,根据用户个体耳图处理音频输入;采用两个自动语音识别(ASR)模块生成音频文本表示,并通过GPT-4o预测两个评分,最终取平均得到估计分数。实验表明,GPT-Whisper-HA相比GPT-Whisper在相对均方根误差(RMSE)上改善2.59%,验证了大模型在零样本语音可懂度预测方面的潜力。

原文摘要 · Abstract (English)

This work focuses on zero-shot non-intrusive speech assessment for hearing aids (HA) using large language models (LLMs). Specifically, we introduce GPT-Whisper-HA, an extension of GPT-Whisper, a zero-shot non-intrusive speech assessment model based on LLMs. GPT-Whisper-HA is designed for speech assessment for HA, incorporating MSBG hearing loss and NAL-R simulations to process audio input based on each individual's audiogram, two automatic speech recognition (ASR) modules for audio-to-text representation, and GPT-4o to predict two corresponding scores, followed by score averaging for the final estimated score. Experimental results indicate that GPT-Whisper-HA achieves a 2.59% relative root mean square error (RMSE) improvement over GPT-Whisper, confirming the potential of LLMs for zero-shot speech assessment in predicting subjective intelligibility for HA users.

助听器大模型语音评估零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。