对比人类与大音视频模型如何理解说话人信息,发现模型尚难模拟人类的社交语言处理机制。
Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment
- 用语义突变和熵值分析模型对说话人与内容矛盾的敏感度
- Qwen2-Audio的突变值能预测人类脑电反应,而Ultravox表现较弱
- 模型无法区分社会规范与生物常识冲突,缺乏人类认知差异
基于语音的AI在处理语言与副语言信息方面面临独特挑战。本研究比较大型音视频模型(LALMs)与人类在语音理解中整合说话者特征的方式,探讨模型是否具备类人认知机制。通过对比Qwen2-Audio与Ultravox 0.5的处理模式与人类脑电图(EEG)响应,利用突变率与熵值指标分析其对社会刻板印象违背(如男性声称常做美甲)及生物知识违背(如男性声称怀孕)的敏感性。结果表明,Qwen2-Audio对说话人-内容不一致表现出更高突变值,且其突变值显著预测人类N400反应;而Ultravox 0.5对说话人特征敏感度有限。更重要的是,两类模型均未复现人类对社会违规(诱发N400)与生物违规(诱发P600)的认知区分。这些发现揭示了当前LALMs在处理说话人上下文语言方面的潜力与局限,提示人类与模型在社交语言处理机制上存在本质差异。
原文摘要 · Abstract (English)
Voice-based AI development faces unique challenges in processing both linguistic and paralinguistic information. This study compares how large audio-language models (LALMs) and humans integrate speaker characteristics during speech comprehension, asking whether LALMs process speaker-contextualized language in ways that parallel human cognitive mechanisms. We compared two LALMs' (Qwen2-Audio and Ultravox 0.5) processing patterns with human EEG responses. Using surprisal and entropy metrics from the models, we analyzed their sensitivity to speaker-content incongruency across social stereotype violations (e.g., a man claiming to regularly get manicures) and biological knowledge violations (e.g., a man claiming to be pregnant). Results revealed that Qwen2-Audio exhibited increased surprisal for speaker-incongruent content and its surprisal values significantly predicted human N400 responses, while Ultravox 0.5 showed limited sensitivity to speaker characteristics. Importantly, neither model replicated the human-like processing distinction between social violations (eliciting N400 effects) and biological violations (eliciting P600 effects). These findings reveal both the potential and limitations of current LALMs in processing speaker-contextualized language, and suggest differences in social-linguistic processing mechanisms between humans and LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。