提出可互操作的索引共享方案,提升信息检索研究的可复现性与协作效率。
Artifact Sharing for Information Retrieval Research
- 设计通用接口规范,支持索引等非代码类资源跨平台共享。
- 通过演示系统验证方案可行性,显著提升索引的发现与使用率。
- 适合需要高效复用检索基础设施的研究者或团队使用。
共享研究产物(如训练好的模型、预构建索引和使用代码)有助于验证中间步骤并提升研究可持续性。尽管代码通常通过 git 仓库、模型通过 HuggingFace Hub 共享已有共识,但索引等其他类型产物仍缺乏统一共享方式。当前研究者多依赖自托管或临时文件传输,限制了资源的可发现性和重用。本演示提出一种灵活且可互操作的信息检索研究产物共享方法,显著提升索引等资源的可访问性与可用性。
原文摘要 · Abstract (English)
Sharing artifacts -- such as trained models, pre-built indexes, and the code to use them -- aids in reproducibility efforts by allowing researchers to validate intermediate steps and improves the sustainability of research by allowing multiple groups to build off one another's prior computational work. Although there are de facto consensuses on how to share research code (through a git repository linked to from publications) and trained models (via HuggingFace Hub), there is no consensus for other types of artifacts, such as built indexes. Given the practical utility of using shared indexes, researchers have resorted to self-hosting these resources or performing ad hoc file transfers upon request, ultimately limiting the artifacts' discoverability and reuse. This demonstration introduces a flexible and interoperable way to share artifacts for Information Retrieval research, improving both their accessibility and usability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。