用语音非语言特征判断共情对话中何时回应最恰当。
Paralinguistic Emotion-Aware Validation Timing Detection in Japanese Empathetic Spoken Dialogue
- 融合语音情绪与声学特征,不依赖文本做响应时机判断。
- 在TESC数据集上准确率显著优于传统语音基线。
- 适合开发更懂情绪的智能对话机器人或心理支持系统。
情感确认是心理治疗中的沟通技巧,通过识别、理解并明确承认对方的情绪和行为来增强信任关系、减轻负面情绪。为最大化情感支持效果,确认的时机和频率至关重要。本研究从语音角度探究确认时机检测方法。基于声学特征与情绪信息,提出一种无需文本上下文的声学-情绪感知模型。具体而言,首先对不同HuBERT骨干网络进行持续自监督训练与微调,获得(i)声学感知自监督学习编码器和(ii)多任务语音情绪分类编码器;随后融合两个编码器,并在下游确认时机检测任务上进一步微调。在TUT情感讲述语料库(TESC)上的实验对比了多种模型、融合机制与训练策略,结果表明所提方法显著优于传统语音基线。研究发现,当结合情绪表征时,非语言语音线索已具备足够信号以判断确认表达的最佳时机,为实现更共情的人机交互提供语音优先路径。
原文摘要 · Abstract (English)
Emotional Validation is a psychotherapy communication technique that involves recognizing, understanding, and explicitly acknowledging another person's feelings and actions, which strengthens alliance and reduces negative affect. To maximize the emotional support provided by validation, it is crucial to deliver it with appropriate timing and frequency. This study investigates validation timing detection from the speech perspective. Leveraging both paralinguistic and emotional information, we propose a paralinguistic- and emotion-aware model for validation timing detection without relying on textual context. Specifically, we first conduct continued self-supervised training and fine-tuning on different HuBERT backbones to obtain (i) a paralinguistics-aware Self-Supervised Learning (SSL) encoder and (ii) a multi-task speech emotion classification encoder. We then fuse these encoders and further fine-tune the combined model on the downstream validation timing detection task. Experimental evaluations on the TUT Emotional Storytelling Corpus (TESC) compare multiple models, fusion mechanisms, and training strategies, and demonstrate that the proposed approach achieves significant improvements over conventional speech baselines. Our results indicate that non-linguistic speech cues, when integrated with affect-related representations, carry sufficient signal to decide when validation should be expressed, offering a speech-first pathway toward more empathetic human-robot interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。