arXiv:2607.02387cs.IRcs.LG2026-07

用智能体搜索解决遥感数据难找问题,提升发现效率。

Bringing Agentic Search to Earth Observation Data Discovery

论文配图:Bringing Agentic Search to Earth Observation Data Discovery
图 1 · 摘自论文原文
  • 基于知识图谱构建智能体搜索系统,支持自然语言查询
  • 新基准测试中召回率提升超5倍,模型性能显著超越传统方法
  • 零样本推理阶段无需训练即可提升28%效果,适合科研人员快速定位数据

NASA及其数据中心拥有数千个地球科学数据集和工具(如Worldview、Giovanni、Science Discovery Engine、Harmony)。即使领域专家也难以找到所需资源。本文提出一个公开部署的智能体搜索系统,可将自然语言研究问题转化为匹配的数据集与工具。我们证明,在大语言模型时代,知识图谱的潜在价值可通过智能体搜索大幅释放。基于NASA地球观测知识图谱(NASA EO-KG),我们构建了包含47,000个查询-数据集对(其中21,000个为任务导向查询)的开放基准NASA-EO-Bench。在该基准上微调的神经评分器优于余弦与BM25基线;进一步结合BM25进行分数融合后,召回率@10(R@10)与平均倒数排名(MRR)均提升超过5倍。在此监督流水线基础上,新增零样本智能体重排阶段,无需额外训练即在分层抽取的N=200子集上使MRR提升28%,表明大模型推理能力与监督检索具有互补性。

原文摘要 · Abstract (English)

NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts. We present an agentic search system, deployed as a public service for the geoscience community, that takes a natural-language research query and returns the matching datasets and tools. We demonstrate that, in the era of large language models, the latent value of knowledge graphs (KGs) can be substantially amplified through agentic search. From the NASA Earth Observation Knowledge Graph (NASA EO-KG) we derive NASA-EO-Bench, an open benchmark of 47k query-dataset pairs (21k task-based queries). A neural scorer fine-tuned on NASA-EO-Bench beats cosine and BM25 baselines. Further combining it with BM25 via score fusion raises both Recall@10 (R@10) and MRR by over 5x. On top of this supervised pipeline, we add a zero-shot agentic reranking stage that, without any additional training, lifts MRR by 28% on a stratified N=200 subset, showing that LLM reasoning is complementary to supervised retrieval.

智能体搜索遥感数据知识图谱大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。