arXiv:2506.03959cs.SDeess.AS2025-06被引 2

用神经信号重建语音,模拟人工耳蜗听觉效果

From Spikes to Speech: NeuroVoc -- A Biologically Plausible Vocoder Framework for Auditory Perception and Cochlear Implant Simulation

  • 通过逆傅里叶变换从神经活动图重建波形
  • 人工耳蜗模型比正常听力模型语音清晰度差7.1 dB
  • 适合研究听觉感知与人工耳蜗仿真

我们提出 NeuroVoc,一种灵活的、模型无关的声码器框架,可从模拟神经活动模式中重构声学波形,采用逆傅里叶变换实现。系统对听觉神经纤维模型输出的时间-频率分箱结果进行简单信号处理。该架构模块化,便于替换或修改底层听觉模型。这避免了在模拟人工耳蜗(CI)用户听觉感知时需为不同编码策略定制声码器的需要,也支持正常听力(NH)与电听觉(EH)模型的直接对比。结果显示,NH模型更完整保留谐波结构。通过在线数字噪声测试(DIN)评估感知可懂度,三种条件:标准语音,以及使用NH和EH模型生成的声码语音。标准DIN与EH声码组的结果分别与临床报告的NH和CI受试者数据无显著差异。平均而言,相比标准测试,NH和EH声码组的言语识别阈(SRT)分别恶化2.4 dB和7.1 dB。表明尽管有退化,该声码器仍能重建可懂语音,并准确反映CI用户在噪声中的性能下降。

原文摘要 · Abstract (English)

We present NeuroVoc, a flexible model-agnostic vocoder framework that reconstructs acoustic waveforms from simulated neural activity patterns using an inverse Fourier transform. The system applies straightforward signal processing to neurogram representations, time-frequency binned outputs from auditory nerve fiber models. Crucially, the model architecture is modular, allowing for easy substitution or modification of the underlying auditory models. This flexibility eliminates the need for speech-coding-strategy-specific vocoder implementations when simulating auditory perception in cochlear implant (CI) users. It also allows direct comparisons between normal hearing (NH) and electrical hearing (EH) models, as demonstrated in this study. The vocoder preserves distinctive features of each model; for example, the NH model retains harmonic structure more faithfully than the EH model. We evaluated perceptual intelligibility in noise using an online Digits-in-Noise (DIN) test, where participants completed three test conditions: one with standard speech, and two with vocoded speech using the NH and EH models. Both the standard DIN test and the EH-vocoded groups were statistically equivalent to clinically reported data for NH and CI listeners. On average, the NH and EH vocoded groups increased SRT compared to the standard test by 2.4 dB and 7.1 dB, respectively. These findings show that, although some degradation occurs, the vocoder can reconstruct intelligible speech under both hearing models and accurately reflects the reduced speech-in-noise performance experienced by CI users.

声码器人工耳蜗听觉模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。