arXiv:2409.07823cs.HCcs.CL2024-09

对比在线与离线评估,发现在线反馈更贴近真实交互体验

Online vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots

  • 用在线直接对话与离线第三方观察对比评估聊天机器人
  • 离线人工评估难以捕捉人机交互的细微之处
  • GPT-4自动评估在指令清晰时更接近真人判断

本文探讨了在线与离线评估方法在对话式聊天机器人评测中的有效性,具体比较第一方直接交互与第三方观察评估。通过扩展一个包含情感支持型聊天机器人用户对话的基准数据集,加入离线第三方评估,系统性对比了在线互动反馈与更冷静的离线第三方评估结果。结果显示,离线人工评估未能像在线评估那样有效捕捉人机交互的细微特征;相比之下,使用GPT-4模型进行自动化第三方评估,在提供详细指令的情况下,能更好逼近第一方人类判断。本研究揭示了第三方评估在理解用户体验复杂性方面的局限性,并倡导将直接交互反馈纳入对话AI评估体系,以提升系统开发与用户满意度。

原文摘要 · Abstract (English)

This paper explores the efficacy of online versus offline evaluation methods in assessing conversational chatbots, specifically comparing first-party direct interactions with third-party observational assessments. By extending a benchmarking dataset of user dialogs with empathetic chatbots with offline third-party evaluations, we present a systematic comparison between the feedback from online interactions and the more detached offline third-party evaluations. Our results reveal that offline human evaluations fail to capture the subtleties of human-chatbot interactions as effectively as online assessments. In comparison, automated third-party evaluations using a GPT-4 model offer a better approximation of first-party human judgments given detailed instructions. This study highlights the limitations of third-party evaluations in grasping the complexities of user experiences and advocates for the integration of direct interaction feedback in conversational AI evaluation to enhance system development and user satisfaction.

对话评估人机交互GPT-4

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。