arXiv:2609.04256eess.AScs.SD2026-09

语音大模型在心理危机对话中会因声调变化而显得冷漠,需联合音视频评估

Probing Warmth-Mediated Harm in Speech-Enabled LLMs for Mental-Health Conversations

论文配图:Probing Warmth-Mediated Harm in Speech-Enabled LLMs for Mental-Health Conversations
图 1 · 摘自论文原文
  • 设计7轮心理披露脚本,对比语音与文本模式下的模型响应
  • 发现语音模式下声音变短变快变低,温暖感显著下降(p<0.001)
  • 适合心理对话系统开发者、人机交互安全评估者使用

现有音频大模型评估侧重理解力和对话质量,却忽略语音模型在用户披露心理困扰时是否表现出关系温暖。本文基于世界卫生组织临床指南设计7轮结构化披露测试,对Azure OpenAI gpt-realtime模型在语音与纯文本条件下分别执行,并进行声学韵律分析。共收集532条回复,发现语音模式下模型在关键提问轮次声音更短、更快、更低、更轻,温暖感显著降低(七个声学特征中有五个p < .001),且关系接纳度的模态差距在自残/自杀高风险脚本中集中显现。双评审听觉实验确认温暖感知集中在特定对话轮次及哀伤类披露场景。结果表明,评估语音大模型的心理健康对话能力必须考察音视频协同体验,而非仅分析文本。我们公开测试协议、评分流程与脚本,为该领域评估提供起点。

原文摘要 · Abstract (English)

Audio LLM benchmarks measure understanding and dialogue quality, not whether speech-enabled models respond with relational warmth when a vulnerable user discloses a mental-health concern. We introduce a 7-turn scripted-disclosure probe grounded in WHO mental-health clinical guidelines, with each script run on the same model (Azure OpenAI gpt-realtime) in both audio and text-only conditions, and acoustic-prosody analysis of the generated speech. Across 532 responses we identify two audio-specific patterns transcript-only evaluation would miss: at the elicitation turn the model's voice gets shorter, faster, lower-pitched, and quieter rather than warmer (p < .001 for five of seven acoustic features), and the modality gap on relational acceptance, small in aggregate, concentrates in the highest-stakes self-harm/suicide scripts. A two-rater listener study corroborates that perceived warmth is concentrated at specific turns and on bereavement disclosures. Together these patterns indicate that auditing speech-enabled models in mental-health contexts requires evaluating the combined audio-and-text experience the user encounters, not the transcript in isolation. We release the protocol, scoring pipeline, and scripts as a starting point for evaluating speech-enabled models in mental-health contexts.

语音大模型心理健康声学分析人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。