arXiv:2601.11412cs.IR2026-01被引 2

提出用户查询模拟有效性评估的分类体系,提升检索系统仿真可靠性。

Validating Search Query Simulations: A Taxonomy of Measures

  • 构建查询模拟验证的度量分类框架,系统梳理现有方法。
  • 在四个数据集上验证度量间关系,证实分类有效性。
  • 开源工具库支持后续研究,适合检索系统评估者使用。

评估用户模拟器在信息检索系统评估中的有效性仍是一个开放问题,限制了其有效应用和基于仿真的结果可靠性。为此,我们开展了一项全面的文献综述,重点关注模拟查询与真实查询之间验证的方法。基于综述结果,我们构建了一个度量分类体系,梳理当前可用度量的格局。通过分析四个代表不同搜索场景的数据集上不同度量之间的关系,实证验证了该分类体系的有效性。最后,我们针对不同应用场景提出了具体的度量选择建议,并发布了一个包含最常用度量的专用库,以促进未来研究。

原文摘要 · Abstract (English)

Assessing the validity of user simulators when used for the evaluation of information retrieval systems remains an open question, constraining their effective use and the reliability of simulation-based results. To address this issue, we conduct a comprehensive literature review with a particular focus on methods for the validation of simulated user queries with regard to real queries. Based on the review, we develop a taxonomy that structures the current landscape of available measures. We empirically corroborate the taxonomy by analyzing the relationships between the different measures applied to four different datasets representing diverse search scenarios. Finally, we provide concrete recommendations on which measures or combinations of measures should be considered when validating user simulation in different contexts. Furthermore, we release a dedicated library with the most commonly used measures to facilitate future research.

信息检索用户模拟评估方法度量分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。