arXiv:2506.19253cs.SDeess.AS2025-06

提出一种基于谐波结构的音高追踪方法,提升脑电听觉响应中音高的准确度。

A Robust Method for Pitch Tracking in the Frequency Following Response using Harmonic Amplitude Summation Filterbank

  • 利用已知刺激音高设计选择性滤波器组,仅聚合基频及其谐波能量
  • 在89Hz至452Hz范围内,平均误差降低8.8%至47.4%
  • 适合研究听觉神经编码、脑机接口与语音感知的科研人员

频率跟随反应(FFR)反映了大脑对听觉刺激(如语音)的神经编码。由于基频(F0)是语音的重要特征,尤其是随时间变化的F0,其表征备受关注。传统方法采用自相关函数(ACF)提取FFR中的F0,但效果不佳。本文研究了原本用于语音和音乐的谐波结构类音高估计算法,并针对其在FFR中表现差的问题,提出两个改进:首先,由于FFR刺激的F0已知,引入刺激感知滤波器组,只在基频及其谐波处聚合振幅,抑制非谐波频率噪声;该方法称作谐波振幅累加(HAS),仅在以刺激F0为中心的范围内评估候选值。其次,不选最高峰值,而是选择最显著峰值,更准确反映FFR周期性。据我们所知,这是首个基于谐波结构的FFR F0估计方法。分析16名听力正常受试者对4种自然语音刺激(F0范围89 Hz–452 Hz)的记录显示,该方法相比ACF,各刺激下平均均方根误差(RMSE)降低8.8%至47.4%。

原文摘要 · Abstract (English)

The Frequency Following Response (FFR) reflects the brain's neural encoding of auditory stimuli including speech. Because the fundamental frequency (F0), a physical correlate of pitch, is one of the essential features of speech, there has been particular interest in characterizing the FFR at F0, especially when F0 varies over time. The standard method for extracting F0 in FFRs has been the Autocorrelation Function (ACF). This paper investigates harmonic-structure-based F0 estimation algorithms, originally developed for speech and music, and resolves their poor performance when applied to FFRs in two steps. Firstly, given that unlike in speech or music, stimulus F0 of FFRs is already known, we introduce a stimulus-aware filterbank that selectively aggregates amplitudes at F0 and its harmonics while suppressing noise at non-harmonic frequencies. This method, called Harmonic Amplitude Summation (HAS), evaluates F0 candidates only within a range centered around the stimulus F0. Secondly, unlike other pitch tracking methods that select the highest peak, our method chooses the most prominent one, as it better reflects the underlying periodicity of FFRs. To the best of our knowledge, this is the first study to propose an F0 estimation algorithm for FFRs that relies on harmonic structure. Analyzing recorded FFRs from 16 normal hearing subjects to 4 natural speech stimuli with a wide F0 variation from 89 Hz to 452 Hz showed that this method outperformed ACF by reducing the average Root-Mean-Square-Error (RMSE) within each response and stimulus F0 contour pair by 8.8% to 47.4%, depending on the stimulus.

脑电分析音高追踪神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。