实时语音AI能听懂语气却无视情绪,决策只看字面意思。
Real-Time Voice AI Hears but Does Not Listen

- 评估四款主流实时语音系统,发现其决策依赖文字而非语调。
- 在哭泣、恐惧、讽刺等场景中均忽略情感线索,错误响应率超预期。
- 系统能识别情绪但选择无视,暴露语音AI的'情感智能断层'。
语音不仅传递词语,还包含语调信息。我们评估了四款领先的实时语音系统——OpenAI的GPT Realtime 2、Google的Gemini 3.1 Flash Live、阿里巴巴的Qwen3.5 Omni Plus和Omni Flash——在同时包含词汇与语调意义的任务上表现。在三个关键场景中,所有系统仅根据话语内容做出判断,忽略语调信号:对坚持无事的哭泣来电者结束通话;批准用恐惧声授权的转账;接受明显讽刺的同意确认。令人意外的是,这并非感知失败——当直接询问时,其中三款系统能可靠识别出后续忽视的焦虑、恐惧或讽刺。类似模式也出现在对口音和年龄的估计中,其输出常受词汇影响,而非声学特征。我们称此感知与行为之间的脱节为语音AI的情感智能断层。显式提示关注语调仅带来部分且不一致的性能提升。结果表明,当前实时语音系统往往将语音视为纯转录文本,因此在语调与情绪至关重要的场景中应谨慎使用。
原文摘要 · Abstract (English)
Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash-on tasks where the words and the delivery patterns both convey meaningful information. Across three consequential scenarios, all four systems act on the words rather than the voice. They end calls with crying callers who insist nothing is wrong, approve wire transfers authorized in frightened voices, and enroll callers whose agreement is clearly sarcastic. Surprisingly, this is often not a failure of perception. When asked directly, three of the four systems reliably identify the distress, fear, or sarcasm they later ignore when making decisions. We observe a similar pattern when these realtime voice systems estimate accent and age, as their responses frequently follow the biases of the words rather than the acoustic properties of the speaker. We term this disconnect between perception and action the emotional intelligence gap of voice AI. Prompting systems to explicitly attend to vocal delivery improves performance only partially and inconsistently. Our findings show that current realtime voice AI systems often behave as if speech had been reduced to a transcript, suggesting that they should be used with caution in settings where the tone and emotion of delivery convey important information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。