arXiv:2602.00352cs.CL2026-02

构建多轮交互式检索评估基准,测试模型在模糊线索下的推理能力

DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning

  • 设计双代理框架,主代理通过多轮提问从记忆代理中检索目标
  • 1011个提示测试显示当前最优模型准确率仅36%(含文本/图像/音频/视频)
  • 适合研究对话系统、多模态推理与记忆增强型AI的开发者使用

在对话中回忆信息时,人们常需经过多轮互动才能想起。然而现有评估基准大多局限于单轮场景。为更真实地模拟这种‘舌尖现象’式的检索过程,我们提出双代理评估基准DETOUR,包含1,011个提示。该基准由主代理(被评估对象)和保持一致的内存代理构成,主代理需通过查询内存代理来识别目标实体。实验结果表明,当前最先进模型在全模态(文本、图像、音频、视频)测试下准确率仅为36%,凸显了在不明确条件下提升检索与推理能力的重要性。

原文摘要 · Abstract (English)

When recalling information in conversation, people often arrive at the recollection after multiple turns. However, existing benchmarks for evaluating agent capabilities in such tip-of-the-tongue search processes are restricted to single-turn settings. To more realistically simulate tip-of-the-tongue search, we introduce Dual-agent based Evaluation Through Obscure Under-specified Retrieval (DETOUR), a dual-agent evaluation benchmark containing 1,011 prompts. The benchmark design involves a Primary Agent, which is the subject of evaluation, tasked with identifying the recollected entity through querying a Memory Agent that is held consistent across evaluations. Our results indicate that current state-of-the-art models still struggle with our benchmark, only achieving 36% accuracy when evaluated on all modalities (text, image, audio, and video), highlighting the importance of enhancing capabilities in underspecified scenarios.

多代理系统对话评估检索推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。