用大模型自动评估对话任务基准的质量优劣
Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

- 不用人工标注,用大模型判断基准一致性、复杂度和覆盖度
- 能有效区分不同质量等级的基准,结果与人工一致
- 适合评估自动生成或人工设计的对话系统评测集
面向任务的对话代理通常通过精心设计或自动生成的基准进行评估,但基准质量很少被检验。低质量基准可能包含不一致的任务、简单的场景或有限的策略覆盖,导致评估不可靠。本文提出一种无需参考的评估框架,利用大模型评判者评估基准的一致性、复杂性和策略覆盖度,并提供可操作的问题诊断。通过与独立人工标注对比验证框架有效性,并对不同能力大模型生成的基准及受控降质处理的基准进行评估。在多个领域和评判模型下,该方法均能稳定区分基准质量层级。进一步证明其适用于人工整理的基准。该框架为合成与人工基准提供了实用的评估路径。
原文摘要 · Abstract (English)
Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework that uses LLM judges to assess benchmark consistency, complexity, and policy coverage, while providing actionable diagnostics of weaknesses. We validate the framework by demonstrating agreement with independent human annotations and by evaluating benchmarks generated by LLMs of varying capabilities, as well as benchmarks subjected to controlled quality-degrading perturbations. Across domains and judge models, the proposed metrics consistently distinguish between benchmark quality levels. We further demonstrate the framework's applicability to manually curated benchmarks. Our framework offers a practical approach for evaluating synthetic and manually curated conversational-agent benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。