ASPIRE 是一个可视化工具,帮助研究人员深入分析信息检索模型的性能差异。
ASPIRE: Assistive System for Performance Evaluation in IR
- 通过交互式界面实现多实验、单查询等多维度对比分析
- 支持查询特征与性能之间的关联探索,揭示模型表现差异原因
- 适合需要细致评估检索系统表现的研究人员使用
信息检索(IR)评估远不止于在表格中展示性能指标。研究者常需在多个维度上比较不同模型的表现,例如精确率-召回率权衡和响应时间,以理解特定查询下各模型表现差异的原因。我们提出 ASPIRE(Assistive System for Performance Evaluation in IR),一个视觉分析工具,旨在通过提供全面且用户友好的界面,解决这些复杂性问题。ASPIRE 支持四个关键方面的 IR 实验评估与分析:单/多实验比较、查询级别分析、查询特征与性能的相互作用分析,以及基于数据集的检索分析。我们以 TREC Clinical Trials 数据集为例展示了 ASPIRE 的功能。ASPIRE 是一个开源工具,可在线获取:https://github.com/GiorgosPeikos/ASPIRE。
原文摘要 · Abstract (English)
Information Retrieval (IR) evaluation involves far more complexity than merely presenting performance measures in a table. Researchers often need to compare multiple models across various dimensions, such as the Precision-Recall trade-off and response time, to understand the reasons behind the varying performance of specific queries for different models. We introduce ASPIRE (Assistive System for Performance Evaluation in IR), a visual analytics tool designed to address these complexities by providing an extensive and user-friendly interface for in-depth analysis of IR experiments. ASPIRE supports four key aspects of IR experiment evaluation and analysis: single/multi-experiment comparisons, query-level analysis, query characteristics-performance interplay, and collection-based retrieval analysis. We showcase the functionality of ASPIRE using the TREC Clinical Trials collection. ASPIRE is an open-source toolkit available online: https://github.com/GiorgosPeikos/ASPIRE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。