构建多轮RAG对话基准,揭示模型在复杂问题上的表现短板。
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
- 设计666个任务覆盖6大领域,含2800+对话轮次
- 模型在无法回答、不明确、非独立问题上错误率超50%
- 适合评估多轮对话系统鲁棒性,研究者必看
我们提出MTRAG-UN,一个用于探索多轮检索增强生成中开放挑战的基准。该基准包含666个任务,覆盖6个领域,共超过2800个对话轮次,并配有配套语料库。实验表明,当前的检索与生成模型在面对无法回答(UNanswerable)、不明确(UNderspecified)、非独立(NONstandalone)的问题以及模糊响应(UNclear)时仍表现不佳。该基准已在GitHub公开:https://github.com/IBM/mt-rag-benchmark。
原文摘要 · Abstract (English)
We present MTRAG-UN, a benchmark for exploring open challenges in multi-turn retrieval augmented generation, a popular use of large language models. We release a benchmark of 666 tasks containing over 2,800 conversation turns across 6 domains with accompanying corpora. Our experiments show that retrieval and generation models continue to struggle on conversations with UNanswerable, UNderspecified, and NONstandalone questions and UNclear responses. Our benchmark is available at https://github.com/IBM/mt-rag-benchmark
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。