arXiv:2506.11112cs.CLcs.HC2025-06

为对话式信息获取系统建立标准化评估框架。

Dagstuhl Perspectives Workshop 24352 -- Conversational Agents: A Framework for Evaluation (CAFE): Manifesto

  • 提出CAFE框架,包含六项核心评估组件。
  • 明确系统目标、用户任务与评估指标的匹配关系。
  • 适合研究对话智能评估的学者与工程师参考。

本次研讨会上,我们深入探讨了对话式信息获取(CONIAC)的定义及其独特特征,提出了抽象化的世界模型,并定义了用于评估CONIAC系统的对话智能评估框架(CAFE)。该框架包含六个主要组成部分:1)系统利益相关方的目标;2)评估中需研究的用户任务;3)执行任务的用户特征;4)需考虑的评估标准;5)采用的评估方法;6)选定的定量指标测量方式。

原文摘要 · Abstract (English)

During the workshop, we deeply discussed what CONversational Information ACcess (CONIAC) is and its unique features, proposing a world model abstracting it, and defined the Conversational Agents Framework for Evaluation (CAFE) for the evaluation of CONIAC systems, consisting of six major components: 1) goals of the system's stakeholders, 2) user tasks to be studied in the evaluation, 3) aspects of the users carrying out the tasks, 4) evaluation criteria to be considered, 5) evaluation methodology to be applied, and 6) measures for the quantitative criteria chosen.

对话系统评估框架信息获取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。