提出可复现的多媒体检索评估框架,降低实验门槛。
Performance Evaluation in Multimedia Retrieval
- 构建形式化模型表达检索实验全要素
- 开发灵活开源评估工具链支持多种场景
- 适合需要可复现实验的多媒体研究者
多媒体检索的性能评估依赖于各类检索实验,涉及多种技术与度量方法。这些实验可能包含人机协同或纯机器流程,且常具高度复杂性与场景特异性,导致难以比较或复现。本文提出一种形式化模型,用于表达检索实验的全部相关要素,并构建一个灵活的开源评估基础设施来实现该模型。该工作旨在降低检索实验的实施门槛,提升实验的可复现性。
原文摘要 · Abstract (English)
Performance evaluation in multimedia retrieval, as in the information retrieval domain at large, relies heavily on retrieval experiments, employing a broad range of techniques and metrics. These can involve human-in-the-loop and machine-only settings for the retrieval process itself and the subsequent verification of results. Such experiments can be elaborate and use-case-specific, which can make them difficult to compare or replicate. In this paper, we present a formal model to express all relevant aspects of such retrieval experiments, as well as a flexible open-source evaluation infrastructure that implements the model. These contributions intend to make a step towards lowering the hurdles for conducting retrieval experiments and improving their reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。