首个面向搜索行为模拟的公开基准,用于预测用户下一步查询。
Sim4IA-Bench: A User Simulation Benchmark Suite for Next Query and Utterance Prediction
- 构建包含160个真实搜索会话的基准数据集,支持多轮模拟预测。
- 提供70个会话中最多62次模拟运行,覆盖查询与语句的下一步预测。
- 适用于信息检索、交互建模研究者,推动可复现的用户模拟研究。
由于缺乏公认评估标准和基准,用户模拟的验证一直面临挑战,难以判断模拟器是否真实反映用户行为。作为SIGIR 2025年Sim4IA研讨会的一部分,我们提出Sim4IA-Bench——首个面向信息检索领域下一次查询与语句预测的用户模拟基准套件。该数据集包含来自CORE搜索引擎的160个真实搜索会话,其中70个会话提供最多62次模拟运行,分为任务A(预测下一次查询)和任务B(预测下一次语句)。该套件为评估与比较用户模拟方法提供了基础,并支持新有效性度量的开发。尽管规模有限,但它是首个将真实搜索会话与模拟的下一步查询预测相连接的公开基准。除作为下一步查询预测的测试平台外,还可用于探索查询重构行为、意图漂移及交互感知的检索评估。我们还引入了一种新的评估指标。通过开源发布,旨在促进可复现研究,推动信息获取领域更真实、可解释的用户模拟发展:https://github.com/irgroup/Sim4IA-Bench。
原文摘要 · Abstract (English)
Validating user simulation is a difficult task due to the lack of established measures and benchmarks, which makes it challenging to assess whether a simulator accurately reflects real user behavior. As part of the Sim4IA Micro-Shared Task at the Sim4IA Workshop, SIGIR 2025, we present Sim4IA-Bench, a simulation benchmark suit for the prediction of the next queries and utterances, the first of its kind in the IR community. Our dataset as part of the suite comprises 160 real-world search sessions from the CORE search engine. For 70 of these sessions, up to 62 simulator runs are available, divided into Task A and Task B, in which different approaches predicted users next search queries or utterances. Sim4IA-Bench provides a basis for evaluating and comparing user simulation approaches and for developing new measures of simulator validity. Although modest in size, the suite represents the first publicly available benchmark that links real search sessions with simulated next-query predictions. In addition to serving as a testbed for next query prediction, it also enables exploratory studies on query reformulation behavior, intent drift, and interaction-aware retrieval evaluation. We also introduce a new measure for evaluating next-query predictions in this task. By making the suite publicly available, we aim to promote reproducible research and stimulate further work on realistic and explainable user simulation for information access: https://github.com/irgroup/Sim4IA-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。