arXiv:2604.27878cs.IR2026-04被引 1

提出评估用户模拟器的统一工具包,区分真实行为与系统评测可靠性

SimEval-IR: A Unified Toolkit and Benchmark Suite for Evaluating User Simulators and Search Sessions

论文配图:SimEval-IR: A Unified Toolkit and Benchmark Suite for Evaluating User Simulators and Search Sessions
图 1 · 摘自论文原文
  • 构建统一会话结构与可复现的评估基准
  • 发现人类相似性检测对系统排序无效,会话嵌入距离更相关
  • 适合交互式检索研究者和评测框架开发者

用户模拟器在交互式信息检索中日益重要,但缺乏标准化评估工具。模拟器需兼顾行为真实性和测试可靠性,二者常被混淆却可能冲突。本文提出 SimEval-IR,一个开源工具包与基准套件,明确区分并量化这两项指标。其包含:(1) 统一会话模式,兼容搜索与对话交互,支持验证数据适配器与显式损失统计;(2) 三个可执行基准,分别评估行为真实性、基于 RATE 风格的测试可靠性,以及两者关联性分析;(3) 在四个真实数据集(两种语言,四类模拟器)上的基线结果。关键发现:主流的分类器-判别器‘人类相似性’检验对系统排名有效性几乎无预测力(r=+0.09, n=48),而点击深度距离与会话嵌入的 Fréchet 距离具有更强信号(|r|=0.43 与 0.40,p≤0.005)。所有配置与脚本均已公开,支持复现。

原文摘要 · Abstract (English)

User simulators are increasingly central to interactive information retrieval, yet the community lacks standardized evaluation tools. Simulators serve two objectives, behavioral realism (matching real user behavior) and tester reliability (producing valid system rankings), and these are often conflated despite being distinct and sometimes conflicting. We present SimEval-IR, an open-source toolkit and benchmark suite that makes this distinction measurable. SimEval-IR provides: (1) a canonical session schema unifying session search and conversational interactions, with validated dataset adapters and explicit loss accounting; (2) three executable benchmarks covering behavioral realism, tester reliability with RATE-style estimation, and an analysis linking the two; and (3) baseline results across four real datasets in two languages and four simulator families. Our key finding: the classifier-discriminator ''human-likeness'' check, the dominant realism test in the literature, has essentially no pooled predictive power for system-ranking validity ($r{=}{+}0.09$, $n{=}48$), while marginal click-depth distance and Fréchet distance over session embeddings give a much stronger signal ($|r|{=}0.43$ and $0.40$, $p{\leq}0.005$). SimEval-IR is released with all configurations and scripts to reproduce the reported analysis.

信息检索用户模拟评估基准可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。