对比人与机器在多语言语音理解中的表现,发现人类母语优势明显,而机器在并行处理上更强。
Benchmarking Humans and Machines on Complex Multilingual Speech Understanding Tasks
- 设计多语言语音问答任务,对比人类与机器在清晰与混合语音下的表现
- 母语者专注目标说话人能力显著优于第二语言,而大模型在单声道中超越人类
- 人类依赖母语化注意力线索,机器则偏好并行信息提取,适合研究语音认知与模型泛化
听觉注意和选择性相位锁定是人类在复杂声学场景及鸡尾酒会环境中理解语音的核心机制,但多语言情境下的相关能力仍不明确。近年来机器对自然语音的理解取得进展,但对重叠与多通道语音的解析仍存疑问。本文提出一种系统化范式,用于研究人类与机器在多语言环境下,针对清晰与混合通道语音的语音问答任务表现。结果显示,人类听者在母语(L1)中对目标说话人的选择性注意显著优于第二语言(L2);而基于语音的大语言模型(LLMs)在单说话人、清晰语音条件下表现匹配或超越人类,但在双说话人场景中常难以实现有效选择性注意。结果揭示关键差异:人类在母语中注意力机制更高效,而模型默认采用并行信息提取,性能超越人类。
原文摘要 · Abstract (English)
Auditory attention and selective phase-locking are central to human speech understanding in complex acoustic scenes and cocktail party settings, yet these capabilities in multilingual subjects remain poorly understood. While machine understanding of natural speech has advanced in recent years, questions persist about comprehension of overlapped and mixed-channel speech. We propose a systematic paradigm for studying humans and machines in speech question-answering tasks in multilingual settings with clean and mixed-channel speech. For human listeners, selective attention to a target speaker was significantly better in their native language (L1) than in their second language (L2). For machine listening, speech-based large language models (LLMs) match or exceed human performance in clean, single-speaker conditions but often struggle to selectively attend in two-speaker settings. These results reveal a key divergence: humans rely on attentional cues that are more streamlined in their native language, whereas LLMs default to parallel information extraction which exceed human skills.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。