arXiv:2409.06097cs.CL2024-09被引 19

构建双语对话评估框架,测试智能体澄清与提问能力。

ClarQ-LLM: A Benchmark for Models Clarifying and Requesting Information in Task-Oriented Dialog

  • 设计31类任务、每类10个对话场景,模拟真实信息获取过程。
  • 引入对话提供方代理,使智能体需主动提问才能完成任务。
  • 大模型成功率仅60.05%,凸显挑战性,适合对话系统研究者使用。

我们提出ClarQ-LLM,一个包含中英双语对话任务、对话代理和评估指标的评测框架,旨在作为评估任务导向对话中智能体提出澄清问题能力的强基准。该框架包含31种不同任务类型,每类有10个独特的信息寻求者与提供者之间的对话场景,要求寻求者通过提问解决不确定性并收集必要信息以完成任务。不同于传统基于固定对话内容的基准,ClarQ-LLM引入了提供者对话代理,复现原始人类提供者的交互行为,使当前及未来寻求者代理能通过与该提供者代理直接互动来测试其信息获取能力。测试结果显示,LLAMA3.1 405B寻求者代理最高成功率为60.05%,表明ClarQ-LLM对后续研究构成严峻挑战。

原文摘要 · Abstract (English)

We introduce ClarQ-LLM, an evaluation framework consisting of bilingual English-Chinese conversation tasks, conversational agents and evaluation metrics, designed to serve as a strong benchmark for assessing agents' ability to ask clarification questions in task-oriented dialogues. The benchmark includes 31 different task types, each with 10 unique dialogue scenarios between information seeker and provider agents. The scenarios require the seeker to ask questions to resolve uncertainty and gather necessary information to complete tasks. Unlike traditional benchmarks that evaluate agents based on fixed dialogue content, ClarQ-LLM includes a provider conversational agent to replicate the original human provider in the benchmark. This allows both current and future seeker agents to test their ability to complete information gathering tasks through dialogue by directly interacting with our provider agent. In tests, LLAMA3.1 405B seeker agent managed a maximum success rate of only 60.05\%, showing that ClarQ-LLM presents a strong challenge for future research.

对话系统信息获取评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。