arXiv:2606.00851cs.SDcs.CL2026-06

让语音助手根据情绪实时调整回应,更懂用户心情。

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning

论文配图:Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning
图 1 · 摘自论文原文
  • 用连续情绪信号控制对话生成,结合语音与多模态感知
  • 在12种情绪锚点下生成语义与情感都匹配的回应
  • 适合做情感交互、智能客服等需要共情的应用

情感化语音对话系统需推断用户情绪以作出恰当回应,但日常语音常携带微弱、中性或模糊的情绪线索。为此,我们提出 Sympatheia,一种基于用户语音推断情绪、并在有可用时引入连续效价-唤醒(VA)控制信号的语音到语音对话框架。为训练模型,我们构建了包含12个情绪锚点的合成语音对话数据集 Sympatheia-18k,其中包含用于学习情感语音行为的情绪划分,以及一组中性查询与多个情绪条件响应配对的中性划分,以在情绪不明确情况下分离显式情绪控制。实验表明,Sympatheia 在生成语义和语音表达均符合情绪的回应方面优于现有基线。进一步结果表明,同一VA接口可整合来自面部表情、生物信号和文本情绪描述等多种传感模块的情绪估计,在仅靠语音提供有限情绪证据时提升回应一致性。这些结果表明,连续情绪条件化是构建情感自适应语音助手的有效实用路径。

原文摘要 · Abstract (English)

Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues. To address this, we introduce Sympatheia, a speech-to-speech dialogue framework conditioned on affect inferred from the user's speech and, when available, explicit affect specifications provided as a continuous valence--arousal (VA) control signal by a multimodal sensing module or user interface. To train our model, we construct Sympatheia-18k, an emotion-conditioned synthetic spoken dialogue corpus with 12 emotion anchors. This dataset includes an emotional split for learning affective speech behavior, and a neutral split that pairs emotionally neutral queries with multiple emotion-conditioned responses to isolate explicit emotion control in emotionally ambiguous cases. Empirical results show that Sympatheia outperforms speech conversational baselines in generating responses whose semantic content and spoken delivery are both emotionally appropriate. We further show that the same VA interface can integrate emotion estimates from diverse sensing modules, including facial expression, biosignals, and textual affect descriptions, improving response alignment when speech alone provides limited emotional evidence. These results suggest that continuous affect conditioning is an effective practical step for building emotionally adaptive voice assistants.

语音对话情感计算多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。