对比五种对话风格的辅导型聊天机器人,发现核心功能比风格更重要。
Substance over Style: Evaluating Proactive Conversational Coaching Agents
- 设计五种不同风格的多轮对话辅导机器人,支持混合发起对话。
- 155次用户实验显示,用户更看重核心功能而非对话风格。
- 专家和大模型评估与用户感受存在显著偏差,提示评估方法需改进。
尽管自然语言处理在对话任务上取得进展,但多数方法聚焦于单轮回应且目标明确。相比之下,辅导场景具有初始目标不明确、通过多轮互动演化、评估标准主观、混合发起对话等特点。本文描述并实现五种具有不同对话风格的多轮辅导机器人,并通过用户研究收集了155次对话的一手反馈。结果表明,用户高度重视核心功能,若缺乏核心功能,仅靠风格会受到负面评价。通过将用户反馈与健康专家及语言模型的第三方评估对比,发现三类评估方式间存在显著差异。研究为对话式辅导机器人的设计与评估提供了洞见,有助于提升以用户为中心的NLP应用。
原文摘要 · Abstract (English)
While NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coaching presents unique challenges with initially undefined goals that evolve through multi-turn interactions, subjective evaluation criteria, mixed-initiative dialogue. In this work, we describe and implement five multi-turn coaching agents that exhibit distinct conversational styles, and evaluate them through a user study, collecting first-person feedback on 155 conversations. We find that users highly value core functionality, and that stylistic components in absence of core components are viewed negatively. By comparing user feedback with third-person evaluations from health experts and an LM, we reveal significant misalignment across evaluation approaches. Our findings provide insights into design and evaluation of conversational coaching agents and contribute toward improving human-centered NLP applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。