arXiv:2608.16686cs.HCcs.CL2026-08中稿 · presentation at AP…

让机器人同时感知用户情绪变化和自身情绪,实现双向共情互动。

Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots

论文配图:Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots
图 1 · 摘自论文原文
  • 双流建模:追踪用户语言与表情情绪变化,同时感知机器人自身情绪状态。
  • 用户评分更高:在5人试点中,共情响应与满意度显著优于传统单向系统。
  • 闭环共情:机器人生成情绪一致的语音与动作,促进双方情绪同步恢复。

具备共情能力的社会机器人应不仅回应用户说了什么,还需捕捉其情绪在交互中的动态演变。然而现有共情对话系统多以文本为中心,将共情简化为从用户情绪到系统回复的一次性映射,难以体现人机间具身的情绪互动。本文提出AffectLoop,一个基于Misty II机器人的多模态、说话者-倾听者情绪动态感知的语音对话系统。该系统实时追踪说话者的语言与面部情绪动态,估计机器人作为倾听者的语言与行为情绪状态,并将双重情绪信号用于指导大语言模型(LLM)生成响应。机器人随后输出简短共情语句及情绪一致的具身行为,形成闭环的说话者-倾听者情绪互动。在包含五名参与者的初步对照实验中,相比仅依赖话语内容的基线系统,本系统获得更高整体评价,尤其在共情响应与用户满意度方面。事后日志分析显示,系统提升了说话者与倾听者间的情绪一致性,并强化了基于效价的应激恢复。初步结果表明,显式建模说话者情绪动态与倾听者情绪状态,可有效提升具身共情交互效果。

原文摘要 · Abstract (English)

Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primarily model empathy as a one-way mapping from the user's emotion to the system response, limiting their ability to capture embodied speaker--listener affective exchange. We present AffectLoop, a multimodal speaker-listener emotion-dynamics-aware spoken dialogue system implemented on the Misty II robot. The system tracks the speaker's verbal and facial affective dynamics, estimates the robot listener's own verbal and behavioral affective state, and conditions LLM-based response generation on both affective streams. The robot then generates a short spoken empathetic response together with emotionally congruent embodied behavior, forming a closed speaker--listener affective loop. We evaluate the system in a pilot within-subject study with five participants, comparing it with an otherwise identical utterance-conditioned baseline that omits the speaker- and listener-affective-state inputs. The proposed system received higher overall impression ratings, especially for empathetic response and user satisfaction. Post-hoc log analysis further showed higher speaker-listener affective alignment and stronger valence-based distress recovery. These preliminary results suggest that explicitly modeling both speaker emotional dynamics and listener affective state can improve embodied empathetic interaction.

共情机器人多模态情绪动态具身交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。