用脑电波识别听者注意力,让语音模型更懂人想听谁。
AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
- 通过脑电数据推断听者专注的说话人
- 在多人对话中生成更符合听者意图的回答
- 适合研究人机交互与感知驱动AI的学者
听觉基础模型(如听觉大语言模型)对所有声音输入一视同仁,但人类听觉具有天然选择性:在复杂声景中会聚焦特定说话人而忽略其他。现有模型未融入这种选择性,限制了其生成与听者意图一致响应的能力。为此,我们提出意图感知的听觉场景理解(II-ASU),并构建听觉注意力驱动的大语言模型(AAD-LLM)。该模型通过颅内脑电图(iEEG)记录解码听者关注的说话人,并据此优化回答生成。先从神经活动预测关注对象,再以该注意力状态条件化输出。我们在多说话人场景下评估了语音描述、转写提取和问答任务,客观与主观评价均显示响应更契合听者意图。本工作首次探索了以听者感知驱动机器听觉的新范式,为未来以用户为中心的听觉系统铺路。演示与代码已公开:https://aad-llm.github.io。
原文摘要 · Abstract (English)
Auditory foundation models, including auditory large language models (LLMs), process all sound inputs equally, independent of listener perception. However, human auditory perception is inherently selective: listeners focus on specific speakers while ignoring others in complex auditory scenes. Existing models do not incorporate this selectivity, limiting their ability to generate perception-aligned responses. To address this, we introduce Intention-Informed Auditory Scene Understanding (II-ASU) and present Auditory Attention-Driven LLM (AAD-LLM), a prototype system that integrates brain signals to infer listener attention. AAD-LLM extends an auditory LLM by incorporating intracranial electroencephalography (iEEG) recordings to decode which speaker a listener is attending to and refine responses accordingly. The model first predicts the attended speaker from neural activity, then conditions response generation on this inferred attentional state. We evaluate AAD-LLM on speaker description, speech transcription and extraction, and question answering in multitalker scenarios, with both objective and subjective ratings showing improved alignment with listener intention. By taking a first step toward intention-aware auditory AI, this work explores a new paradigm where listener perception informs machine listening, paving the way for future listener-centered auditory systems. Demo and code available: https://aad-llm.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。