arXiv:2601.09366cs.IR2026-01

构建可复用的交互式搜索数据集,助力研究用户行为差异。

LISP -- A Rich Interaction Dataset and Loggable Interactive Search Platform

  • 收集61人共122次会话的详细交互日志与用户特征。
  • 涵盖感知速度、兴趣、专业度等变量,支持多因素分析。
  • 开源工具链,支持研究复现与用户模拟器开发。

我们提出一个可复用的数据集及配套基础设施,用于研究交互式信息检索(IIR)中的人类搜索行为。数据集整合了61名参与者(共122次会话)的详细交互日志,以及感知速度、主题兴趣、搜索专业度和人口统计学信息等用户特征。为促进可复现性与再利用,我们提供了完整的研究设置文档、基于Web的感知速度测试工具,以及开展类似用户研究的框架。本工作使研究者能够探究个体与情境因素对搜索行为的影响,并开发或验证考虑此类变异性的用户模拟器。我们通过示例分析展示了数据集的潜力,并将所有资源开放获取,支持IIR领域内的可复现研究与资源共享。

原文摘要 · Abstract (English)

We present a reusable dataset and accompanying infrastructure for studying human search behavior in Interactive Information Retrieval (IIR). The dataset combines detailed interaction logs from 61 participants (122 sessions) with user characteristics, including perceptual speed, topic-specific interest, search expertise, and demographic information. To facilitate reproducibility and reuse, we provide a fully documented study setup, a web-based perceptual speed test, and a framework for conducting similar user studies. Our work allows researchers to investigate individual and contextual factors affecting search behavior, and to develop or validate user simulators that account for such variability. We illustrate the datasets potential through an illustrative analysis and release all resources as open-access, supporting reproducible research and resource sharing in the IIR community.

交互搜索用户行为数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。