arXiv:2607.05443cs.IRcs.AI2026-07

构建首个跨领域科学代码搜索数据集与基准,助力科研工具高效发现。

Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark

  • 收集5264个科学领域代码库,覆盖五大航天科研方向。
  • 设计219条专家级查询,支持跨领域代码检索评测。
  • 提供万级代码片段与查询对,适配多语言科研场景。

科学家日益依赖开源工具支持研究工作,但在超过6亿个GitHub仓库中定位相关软件仍具挑战。现有代码搜索基准主要面向通用软件工程任务,难以反映科学计算领域的专业术语与需求。本文构建了一个包含5,264个高质量、领域分类的科学代码库的语料库,涵盖美国宇航局科学使命局五大部门:地球科学、天体物理学、行星科学、日地物理及生物与物理科学。所有代码库均配有清洗后的README、提取的主题标签及来自爬取链接的附加上下文。基于此语料库,我们提出两个新型信息检索基准:(1)由领域科学家精心设计的219条专家级代码库查询;(2)包含117,950个代码片段与119,720个查询的超大规模代码片段检索基准,覆盖七种编程语言。基线评估显示,不同科学领域间性能差异显著;代码片段检索亦面临挑战,主要源于各科研群体在文档规范、编码标准和语言习惯上的差异。所有数据集与基准已公开发布于HuggingFace,以支持科学工具发现的研究。

原文摘要 · Abstract (English)

Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub repositories remains challenging. Existing code search benchmarks focus on general software engineering tasks and fail to capture the domain-specific vocabulary and needs of scientific computing. We present a curated corpus of 5,264 high-quality, domain-classified scientific repositories spanning five NASA Science Mission Directorate divisions -- Earth Science, Astrophysics, Planetary Science, Heliophysics, and Biological & Physical Sciences -- enriched with cleaned READMEs, extracted topics, and additional context from crawled links. Building on this corpus, we introduce two novel information retrieval benchmarks: (1) a repository search benchmark with 219 expert-curated queries designed by domain scientists, and (2) a large-scale code snippet retrieval benchmark containing 117,950 code snippets and 119,720 queries across seven programming languages. Baseline evaluations on repository search reveal significant performance variation across scientific domains. Code snippet retrieval proves equally challenging, with substantial variation driven by differing documentation practices, coding standards, and programming language conventions across scientific communities. All datasets and benchmarks are publicly released on HuggingFace to support research on scientific tool discovery.

代码搜索科学计算数据集信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。