用AI生成对话数据,低成本构建更真实的对话检索评估基准。
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks

- 用大模型自动检测现有评测集的语义偏差
- 多智能体系统以1/400成本生成高保真对话
- 新基准支持真实场景挑战,适合研究检索增强生成
准确评估对话式检索对提升检索增强生成(RAG)系统至关重要。然而,现有对话检索评测集存在人工标注成本高、数据稀疏或自动化规则僵硬、不自然等问题。为此,我们提出MTR-Suite,一个统一的框架,用于审计、生成和评测检索能力。其包含:(1) MTR-Eval,基于大模型的审计工具,量化现有评测集中的语义对齐差距;(2) MTR-Pipeline,一种多智能体系统,采用贪婪遍历聚类方法,以1/400的人工成本生成高保真对话;(3) MTR-Bench,一个严格的通用领域评测基准,模拟真实生产环境中的挑战(如复杂话题切换、表达冗长),具备更强的区分能力。代码与数据已开源,地址为https://github.com/rangehow/mtr-suite。
原文摘要 · Abstract (English)
Accurate evaluation of conversational retrieval is pivotal for advancing Retrieval-Augmented Generation (RAG) systems. However, existing conversational retrieval benchmarks suffer from costly, sparse human annotation or rigid, unnatural automated heuristics. To address these challenges, we introduce MTR-Suite, a unified framework for auditing, synthesizing, and benchmarking retrieval. It features: (1) MTR-Eval, an LLM-based auditor quantifying alignment gaps in previous benchmarks; (2) MTR-Pipeline, a multi-agent system using greedy traversal clustering to generate high-fidelity dialogues at 1/400th human cost; and (3) MTR-Bench, a rigorous general-domain benchmark. MTR-Bench mimics production-style challenges (hard topic switching, verbosity), offering superior discriminative power. We make our code and data publicly available to facilitate future research at https://github.com/rangehow/mtr-suite.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。