arXiv:2603.09324cs.HCcs.AI2026-03中稿 · CHI EA 2026

让虚拟角色听懂语气中的情绪,回应更自然有温度。

Reading the Mood Behind Words: Integrating Prosody-Derived Emotional Context into Socially Responsive VR Agents

  • 用语音韵律识别情绪,注入对话上下文影响回复风格。
  • 30人实验显示93.3%用户更喜欢带情绪感知的虚拟角色。
  • 适合开发情感交互型虚拟助手或沉浸式社交应用。

在与具身对话代理的虚拟现实互动中,用户的情绪意图往往通过语调而非字面内容传达。然而,多数虚拟现实代理系统依赖语音转文字处理,丢弃了韵律线索,导致语义正确但情感不符的回应。本文提出一种情感上下文感知的虚拟现实交互流程,将语音情绪作为大语言模型对话代理的显式对话上下文。一个实时语音情绪识别模型从韵律中推断用户情绪状态,情绪标签被注入代理对话上下文以塑造回应的语调与风格。一项包含30名参与者的自身对照实验显示,该方法显著提升了对话质量、自然度、参与感、亲密度和拟人感,93.3%的参与者更偏好情绪感知代理。

原文摘要 · Abstract (English)

In VR interactions with embodied conversational agents, users' emotional intent is often conveyed more by how something is said than by what is said. However, most VR agent pipelines rely on speech-to-text processing, discarding prosodic cues and often producing emotionally incongruent responses despite correct semantics. We propose an emotion-context-aware VR interaction pipeline that treats vocal emotion as explicit dialogue context in an LLM-based conversational agent. A real-time speech emotion recognition model infers users' emotional states from prosody, and the resulting emotion labels are injected into the agent's dialogue context to shape response tone and style. Results from a within-subjects VR study (N=30) show significant improvements in dialogue quality, naturalness, engagement, rapport, and human-likeness, with 93.3% of participants preferring the emotion-aware agent.

虚拟现实情绪识别对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。