首个对话数据检索基准,评估产品洞察检索效果。
Finding Diamonds in Conversation Haystacks: A Benchmark for Conversational Data Retrieval
- 构建涵盖5类分析任务的1.6千条查询与9.1千段对话
- 顶尖模型在NDCG@10上仅达0.51,显示能力差距
- 揭示隐式状态、轮次动态等核心挑战,适合研发对话检索系统者
我们提出首个对话数据检索(CDR)基准,用于评估从对话中提取产品洞察信息的系统性能。该基准包含1.6千条查询、覆盖五类分析任务及9.1千段真实对话,为衡量对话数据检索能力提供可靠标准。对16个主流嵌入模型的评估显示,即使最佳模型在NDCG@10上也仅达0.51,反映出文档与对话检索能力间的显著差距。研究识别出对话检索中的独特挑战:隐式状态识别、轮次动态变化与上下文指代理解,并提供实用查询模板与多任务类别误差分析。数据集与代码已开源:https://github.com/l-yohai/CDR-Benchmark。
原文摘要 · Abstract (English)
We present the Conversational Data Retrieval (CDR) benchmark, the first comprehensive test set for evaluating systems that retrieve conversation data for product insights. With 1.6k queries across five analytical tasks and 9.1k conversations, our benchmark provides a reliable standard for measuring conversational data retrieval performance. Our evaluation of 16 popular embedding models shows that even the best models reach only around NDCG@10 of 0.51, revealing a substantial gap between document and conversational data retrieval capabilities. Our work identifies unique challenges in conversational data retrieval (implicit state recognition, turn dynamics, contextual references) while providing practical query templates and detailed error analysis across different task categories. The benchmark dataset and code are available at https://github.com/l-yohai/CDR-Benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。