arXiv:2605.22465cs.CL2026-05

用语音熵模拟大脑记忆缓冲区,区分信息干扰与物理噪声影响

In Silico Modeling of the RAMPHO Buffer: Dissociating Informational and Energetic Masking via Phonetic Entropy in Deep Neural Networks

论文配图:In Silico Modeling of the RAMPHO Buffer: Dissociating Informational and Energetic Masking via Phonetic Entropy in Deep Neural Networks
图 1 · 摘自论文原文
  • 用wav2vec 2.0的语音熵建模大脑记忆缓冲区机制
  • 高信噪比下破坏干扰语义可减轻认知负担,但低信噪比下损害时间线索
  • 揭示听觉认知与声学增强的帕累托权衡,适合听觉模型研究者

在多说话人环境中听觉理解的核心挑战源于认知瓶颈,根据ELU模型,这表现为RAMPHO情景缓冲区的失效。当前语音增强深度神经网络仅优化物理声学特性,忽略了信息掩蔽带来的认知代价。本文通过自监督声学模型wav2vec 2.0的逐帧语音熵,对RAMPHO缓冲区进行体外模拟。在信噪比(SNR)扫描中对比语义完整干扰与相位去相关干扰(浓度屏障),成功分离了信息干扰的认知代价与能量衰减的物理代价。模拟揭示出认知-声学帕累托优化问题:在高SNR下破坏干扰语义可缓解信息掩蔽,但在低SNR下会严重损害时间片段线索。

原文摘要 · Abstract (English)

The fundamental challenge of listening in multi-talker environments is a cognitive bottleneck, defined by the Ease of Language Understanding (ELU) model as a failure within the RAMPHO episodic buffer. Current deep neural networks for speech enhancement optimize purely for physical acoustics, failing to account for the cognitive penalty of informational masking. Here, we present an in silico simulation of the RAMPHO buffer using the frame-by-frame phonetic entropy of a self-supervised acoustic model (wav2vec 2.0). By contrasting a semantically intact distractor with a phase-decorrelated distractor (the Concentration Shield) across a signal-to-noise ratio (SNR) sweep, we successfully dissociate the cognitive penalty of informational distraction from the physical penalty of energetic decay. The simulation reveals a cognitive-acoustic Pareto optimization problem: destroying a distractor's semantic payload provides a release from informational masking at high SNRs, but fundamentally degrades temporal glimpsing cues at low SNRs.

语音增强认知建模语音熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。