arXiv:2412.01401eess.SPcs.SD2024-12被引 5

线性重建法在多模态听觉注意解码数据集上实现高精度,证明旧模型失败是因眼球追踪干扰而非数据问题。

Linear stimulus reconstruction works on the KU Leuven audiovisual, gaze-controlled auditory attention decoding dataset

  • 采用线性刺激重建算法,直接从脑电解码听觉注意力方向。
  • 在每种条件下准确率均显著高于随机水平,且跨条件、跨被试、跨数据集泛化良好。
  • 提供简单基线评估流程,可作为未来算法的最小基准参考。

本文报告了在线性刺激重建方法在KU Leuven多模态、注视控制的听觉注意解码(AV-GC-AAD)数据集上的表现。该数据集记录了参与者在不同视听条件下,选择关注双语者之一时的脑电图(EEG)信号,旨在分离注视方向与听觉注意方向,以揭示现有空间听觉注意解码(AAD)算法中由眼球运动引起的偏差。此前多种基于空间的AAD方法在该数据集上未能达到显著高于随机水平的性能,常被归因于数据量不足、条件异质性强或被试难以集中注意力。然而本研究显示,线性刺激重建算法在各独立条件下均实现高精度解码,并能有效跨条件、跨新被试甚至跨数据集泛化。结果表明,旧有模型失败并非数据集缺陷所致。此外,本文还提供包含源代码的基线评估流程,可作为未来所有在此数据集上评估的AAD算法的最小基准。

原文摘要 · Abstract (English)

In a recent paper, we presented the KU Leuven audiovisual, gaze-controlled auditory attention decoding (AV-GC-AAD) dataset, in which we recorded electroencephalography (EEG) signals of participants attending to one out of two competing speakers under various audiovisual conditions. The main goal of this dataset was to disentangle the direction of gaze from the direction of auditory attention, in order to reveal gaze-related shortcuts in existing spatial AAD algorithms that aim to decode the (direction of) auditory attention directly from the EEG. Various methods based on spatial AAD do not achieve significant above-chance performances on our AV-GC-AAD dataset, indicating that previously reported results were mainly driven by eye gaze confounds in existing datasets. Still, these adverse outcomes are often discarded for reasons that are attributed to the limitations of the AV-GC-AAD dataset, such as the limited amount of data to train a working model, too much data heterogeneity due to different audiovisual conditions, or participants allegedly being unable to focus their auditory attention under the complex instructions. In this paper, we present the results of the linear stimulus reconstruction AAD algorithm and show that high AAD accuracy can be obtained within each individual condition and that the model generalizes across conditions, across new subjects, and even across datasets. Therefore, we eliminate any doubts that the inadequacy of the AV-GC-AAD dataset is the primary reason for the (spatial) AAD algorithms failing to achieve above-chance performance when compared to other datasets. Furthermore, this report provides a simple baseline evaluation procedure (including source code) that can serve as the minimal benchmark for all future AAD algorithms evaluated on this dataset.

听觉注意解码脑电分析多模态数据线性重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。