arXiv:2604.21096cs.IRcs.CL2026-04

构建首个多语言舌尖查询基准,支持中日韩英四语测试。

Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation

  • 用大模型生成跨语言的'舌尖现象'查询,模拟真实用户难题。
  • 每种语言5000条查询,覆盖多个领域,验证了生成质量。
  • 指南:非英文源数据对生成效果关键,英文维基可作补充。

舌尖现象(ToT)检索评测长期局限于英语,限制了其在多语言信息获取中的应用。本文构建了中文、日文、韩文和英文的多语言ToT测试集,采用基于大语言模型的查询生成框架。系统研究提示语语言与源文档语言对模拟查询保真度的影响,并通过系统排名相关性验证合成查询的有效性。结果表明,有效的ToT模拟需要语言感知的设计:非英语源数据通常更为重要;当非英语来源信息不足时,英文维基百科可提供有益补充。基于此,我们发布了四个涵盖多领域的ToT测试集,每种语言包含5000条查询。本工作提供了首个大规模多语言ToT基准,并为构建非英语环境下的真实查询数据集提供了实用指导。

原文摘要 · Abstract (English)

Tip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this work, we construct multilingual ToT test collections for Chinese, Japanese, Korean, and English, using an LLM-based query simulation framework. We systematically study how prompt language and source document language affect the fidelity of simulated ToT queries, validating synthetic queries through system rank correlation against real user queries. Our results show that effective ToT simulation requires language-aware design choices: non-English language sources are generally important, while English Wikipedia can be beneficial when non-English sources provide insufficient information for query generation. Based on these findings, we release four ToT test collections with 5,000 queries per language across multiple domains. This work provides the first large-scale multilingual ToT benchmark and offers practical guidance for constructing realistic ToT datasets beyond English.

多语言检索评测大模型生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。