评估检索系统随时间变化的稳定性,发现效果好不等于持久好。
Replicability Measures for Longitudinal Information Retrieval Evaluation
- 用动态测试集衡量系统长期有效性,考察排名变化
- 系统性能随时间逐步下降,初始效果强的未必持久
- 适合关注系统长期可靠性的评估研究者
信息检索(IR)系统面临文档更新、用户需求变化等持续演化。尽管期望系统保持稳定效用,但传统评测依赖固定实验设置。基于LongEval共享任务与测试集,本研究探索在动态环境中如何评估检索效果的可复现性。重点考察了效果持久性这一复现性问题,发现系统有效性随时间逐步下降。采用改进的可复现性度量方法,揭示出不同时间点和评测指标下系统排名存在显著差异。结论表明:初始表现最优的系统并不一定具有最强的性能持久性。
原文摘要 · Abstract (English)
Information Retrieval (IR) systems are exposed to constant changes in most components. Documents are created, updated, or deleted, the information needs are changing, and even relevance might not be static. While it is generally expected that the IR systems retain a consistent utility for the users, test collection evaluations rely on a fixed experimental setup. Based on the LongEval shared task and test collection, this work explores how the effectiveness measured in evolving experiments can be assessed. Specifically, the persistency of effectiveness is investigated as a replicability task. It is observed how the effectiveness progressively deteriorates over time compared to the initial measurement. Employing adapted replicability measures provides further insight into the persistence of effectiveness. The ranking of systems varies across retrieval measures and time. In conclusion, it was found that the most effective systems are not necessarily the ones with the most persistent performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。