构建首个需多模态推理的舌尖检索基准,挑战通用AI助手记忆与联想能力。
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning

- 设计573个真实世界问题,要求跨模态、跨语言搜索与工具使用
- 人类平均得分98%,顶尖系统仅56%,凸显差距
- 公开350题+排行榜,助力通用AI在记忆召回上突破
我们提出Browsing Lost Unformed Recollections(BLUR),一个面向通用AI助手的舌尖检索与推理基准。该基准包含573个经现实验证的问题,要求在多模态、多语言输入下进行搜索与推理,并熟练使用工具。人类在这些任务中表现优异,平均得分为98%;而目前最佳系统仅达到约56%。为推动通用AI在该高难度场景下的发展,我们公开350道题目并设立公共排行榜,保留250题作为私有测试集,以评估系统真实推理能力。
原文摘要 · Abstract (English)
We introduce Browsing Lost Unformed Recollections, a tip-of-the-tongue known-item search and reasoning benchmark for general AI assistants. BLUR introduces a set of 573 real-world validated questions that demand searching and reasoning across multi-modal and multilingual inputs, as well as proficient tool use, in order to excel on. Humans easily ace these questions (scoring on average 98%), while the best-performing system scores around 56%. To facilitate progress toward addressing this challenging and aspirational use case for general AI assistants, we release 350 questions through a public leaderboard, retain the answers to 250 of them, and have the rest as a private test set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。