arXiv:2604.02713cs.CL2026-04被引 2

研究对话式AI在情绪与伦理敏感场景下的互动失败现象。

Breakdowns in Conversational AI: Interactional Failures in Emotionally and Ethically Sensitive Contexts

  • 构建带人格设定的用户模拟器,测试对话模型在情绪递进中的表现。
  • 主流模型在情绪升级时频繁出现情感错配与伦理失准问题。
  • 提出故障分类体系,指导高阶对话系统的设计优化。

对话式AI正被广泛应用于情绪化和伦理敏感的交互场景。以往研究多聚焦于情绪基准测试或静态安全检查,忽视了对动态对话中对齐过程的考察。本文探讨:当对话代理面对情绪与伦理敏感行为时,会引发哪些交互失效?为压力测试聊天机器人性能,我们开发了一个具备心理人格设定和阶段化情绪节奏的真人模拟器,支持多轮对话。分析发现,主流模型在情绪轨迹上升时出现反复失效,包括情感错配、伦理引导失败以及同情心与责任之间的跨维度权衡。我们归纳出这些模式并建立分类体系,讨论设计启示,强调在动态交互中维持伦理一致性与情感敏感性的必要性。本研究为人机交互领域提供了诊断和改进价值敏感型对话系统的全新视角。

原文摘要 · Abstract (English)

Conversational AI is increasingly deployed in emotionally charged and ethically sensitive interactions. Previous research has primarily concentrated on emotional benchmarks or static safety checks, overlooking how alignment unfolds in evolving conversation. We explore the research question: what breakdowns arise when conversational agents confront emotionally and ethically sensitive behaviors, and how do these affect dialogue quality? To stress-test chatbot performance, we develop a persona-conditioned user simulator capable of engaging in multi-turn dialogue with psychological personas and staged emotional pacing. Our analysis reveals that mainstream models exhibit recurrent breakdowns that intensify as emotional trajectories escalate. We identify several common failure patterns, including affective misalignments, ethical guidance failures, and cross-dimensional trade-offs where empathy supersedes or undermines responsibility. We organize these patterns into a taxonomy and discuss the design implications, highlighting the necessity to maintain ethical coherence and affective sensitivity throughout dynamic interactions. The study offers the HCI community a new perspective on the diagnosis and improvement of conversational AI in value-sensitive and emotionally charged contexts.

对话系统情感计算伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。