arXiv:2511.21909cs.CL2025-11被引 1

用沃兹技术构建情感引导对话数据集,助力客服情绪智能研究。

A Customer Journey in the Land of Oz: Leveraging the Wizard of Oz Technique to Model Emotions in Customer Service Interactions

  • 通过控制性沃兹实验诱导特定情绪轨迹,生成真实对话数据。
  • 中性情绪占主导,感激与渴望是主要非中性情绪,自评与他评差异显著。
  • 适合研究情绪感知、对话策略优化的AI客服开发者与人机交互学者。

情感感知的客户服务需要领域内对话数据、丰富标注和预测能力,但现有资源多为跨领域、标签单一且侧重事后检测。为此,我们开展受控沃兹(WOZ)实验,诱发具有目标情感轨迹的互动。生成的数据集EmoWOZ-CS包含2,148条来自179名参与者的双语(荷兰语-英语)书面对话,涵盖航空、电商、在线旅游和电信场景。贡献有三:(1)评估沃兹驱动的情绪轨迹设计在情感研究中的有效性;(2)量化人工标注表现与变异,包括自我报告与第三方判断的分歧;(3)基准测试实时支持中的情绪检测与前瞻性情绪推断。结果显示,中性情绪占主导;渴望与感激是最常见的非中性情绪。多标签情绪与效价的共识为中等,唤醒度与支配感更低;自我报告与第三方标签差异明显,中性、感激和愤怒最一致。客观策略常引发中性或感激,次优策略则加剧愤怒、烦扰、失望、渴望与困惑。某些情感策略(如愉快、感激)促进积极回应,而道歉、共情也可能引发渴望、愤怒或烦扰。时间分析表明,对话层面的情绪引导成功实现,尤其针对负面目标;正向与中性目标最终效价分布相似。基准测试凸显从先前对话轮次预测未来情绪的难度,揭示主动情绪感知支持的复杂性。

原文摘要 · Abstract (English)

Emotion-aware customer service needs in-domain conversational data, rich annotations, and predictive capabilities, but existing resources for emotion recognition are often out-of-domain, narrowly labeled, and focused on post-hoc detection. To address this, we conducted a controlled Wizard of Oz (WOZ) experiment to elicit interactions with targeted affective trajectories. The resulting corpus, EmoWOZ-CS, contains 2,148 bilingual (Dutch-English) written dialogues from 179 participants across commercial aviation, e-commerce, online travel agencies, and telecommunication scenarios. Our contributions are threefold: (1) Evaluate WOZ-based operator-steered valence trajectories as a design for emotion research; (2) Quantify human annotation performance and variation, including divergences between self-reports and third-party judgments; (3) Benchmark detection and forward-looking emotion inference in real-time support. Findings show neutral dominates participant messages; desire and gratitude are the most frequent non-neutral emotions. Agreement is moderate for multilabel emotions and valence, lower for arousal and dominance; self-reports diverge notably from third-party labels, aligning most for neutral, gratitude, and anger. Objective strategies often elicit neutrality or gratitude, while suboptimal strategies increase anger, annoyance, disappointment, desire, and confusion. Some affective strategies (cheerfulness, gratitude) foster positive reciprocity, whereas others (apology, empathy) can also leave desire, anger, or annoyance. Temporal analysis confirms successful conversation-level steering toward prescribed trajectories, most distinctly for negative targets; positive and neutral targets yield similar final valence distributions. Benchmarks highlight the difficulty of forward-looking emotion inference from prior turns, underscoring the complexity of proactive emotion-aware support.

情感识别对话系统沃兹实验客服智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。