arXiv:2502.17776cs.IRcs.CL2025-02被引 10

用大模型和真人模拟遗忘查询,提升检索系统评估效果。

Tip of the Tongue Query Elicitation for Simulated Evaluation

  • 用大模型生成合成遗忘查询,可大规模替代人工收集。
  • 生成的查询与真实数据相关性高,能有效评估检索系统性能。
  • 适合研究搜索系统在记忆模糊场景下的表现,或做评测基准。

当用户难以回忆特定标识符(如文档标题)时会出现舌尖现象(TOT)搜索。现有搜索系统对此类场景支持不足,而当前研究受限于查询数据收集困难,主要依赖社区问答网站,存在劳动密集和领域偏倚问题。为此,本文提出两种生成TOT查询的方法:基于大语言模型的模拟器可规模化生成合成查询,在电影领域与社区问答数据生成的查询具有高度相关性;同时设计视觉提示界面,通过引导参与者进入舌尖状态,采集自然真实的查询。在电影、地标和人物三个领域验证表明,两类生成方式均能有效模拟真实查询,且语言相似度高。所生成的查询已发布于TREC 2024 TOT赛道,人类采集数据将于TREC 2025纳入。此外,源代码与视觉刺激素材均已公开。

原文摘要 · Abstract (English)

Tip-of-the-tongue (TOT) search occurs when a user struggles to recall a specific identifier, such as a document title. While common, existing search systems often fail to effectively support TOT scenarios. Research on TOT retrieval is further constrained by the challenge of collecting queries, as current approaches rely heavily on community question-answering (CQA) websites, leading to labor-intensive evaluation and domain bias. To overcome these limitations, we introduce two methods for eliciting TOT queries - leveraging large language models (LLMs) and human participants - to facilitate simulated evaluations of TOT retrieval systems. Our LLM-based TOT user simulator generates synthetic TOT queries at scale, achieving high correlations with how CQA-based TOT queries rank TOT retrieval systems when tested in the Movie domain. Additionally, these synthetic queries exhibit high linguistic similarity to CQA-derived queries. For human-elicited queries, we developed an interface that uses visual stimuli to place participants in a TOT state, enabling the collection of natural queries. In the Movie domain, system rank correlation and linguistic similarity analyses confirm that human-elicited queries are both effective and closely resemble CQA-based queries. These approaches reduce reliance on CQA-based data collection while expanding coverage to underrepresented domains, such as Landmark and Person. LLM-elicited queries for the Movie, Landmark, and Person domains have been released as test queries in the TREC 2024 TOT track, with human-elicited queries scheduled for inclusion in the TREC 2025 TOT track. Additionally, we provide source code for synthetic query generation and the human query collection interface, along with curated visual stimuli used for eliciting TOT queries.

信息检索大模型应用用户行为模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。