通过角色互换研究用户如何调节语音助手情绪,提升情感智能交互体验。
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
- 让人类扮演助手角色调节AI情绪,观察真实情感应对策略。
- 中性或正向回应最受青睐,能有效缓解负面情绪。
- 发现音高、语速等声学特征与情绪状态高度相关,适合用于实时识别。
随着语音助手日益融入日常生活,具备识别并适当回应用户情绪的能力成为关键需求。尽管语音情绪识别(SER)和情感分析已取得进展,但对负面情绪的有效应对仍具挑战。本研究采用角色互换方法,让参与者主动调节AI的情绪表现,而非接收预设回应,探索人类在人机交互中的情感应对策略。结合声学特征分析与自然语言处理(NLP),研究考察了不同情绪情境下的语音模式。结果显示,当面对负面情绪线索时,用户更偏好中性或正向回应,体现出自然的情感调节与降级倾向。关键声学指标如均方根(RMS)、零交叉率(ZCR)和抖动(jitter)对情绪状态敏感;而情感极性与词汇多样性(TTR)可有效区分正负向回应。这些发现为构建自适应、情境感知的语音助手提供了重要依据,使其能提供更具同理心、文化敏感性与用户契合度的响应,从而增强人机交互中的信任与参与感。
原文摘要 · Abstract (English)
As voice assistants (VAs) become increasingly integrated into daily life, the need for emotion-aware systems that can recognize and respond appropriately to user emotions has grown. While significant progress has been made in speech emotion recognition (SER) and sentiment analysis, effectively addressing user emotions-particularly negative ones-remains a challenge. This study explores human emotional response strategies in VA interactions using a role-swapping approach, where participants regulate AI emotions rather than receiving pre-programmed responses. Through speech feature analysis and natural language processing (NLP), we examined acoustic and linguistic patterns across various emotional scenarios. Results show that participants favor neutral or positive emotional responses when engaging with negative emotional cues, highlighting a natural tendency toward emotional regulation and de-escalation. Key acoustic indicators such as root mean square (RMS), zero-crossing rate (ZCR), and jitter were identified as sensitive to emotional states, while sentiment polarity and lexical diversity (TTR) distinguished between positive and negative responses. These findings provide valuable insights for developing adaptive, context-aware VAs capable of delivering empathetic, culturally sensitive, and user-aligned responses. By understanding how humans naturally regulate emotions in AI interactions, this research contributes to the design of more intuitive and emotionally intelligent voice assistants, enhancing user trust and engagement in human-AI interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。