用三阶段大模型流水线,从地球观测数据中自动发现新研究假设。
EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation

- 基于知识图谱和图神经网络,筛选潜在的数据组合。
- 在1475个数据集上生成160个跨领域的科学假设。
- 结果可靠且新颖,适合科研人员探索跨学科方向。
大语言模型被用于科学假设生成,但多数工作依赖非结构化文献与自由文本。本文提出一个面向地球观测的流水线,将假设生成直接锚定于NASA地球观测知识图谱。通过训练于历史共使用关系的异构图神经网络,对候选数据集组合进行排序;再由三阶段大模型流程过滤、生成并评估结构化研究假设。该系统应用于1475个NASA数据集,生成了涵盖生态水文、冰川学、气溶胶-云相互作用、植被物候及平流层化学等多领域的160个假设。模型预测的新颖数据组合在可接受性评分上接近真实文献中的共使用案例,表明该流水线能挖掘出科学上合理且未被探索的组合。2×2×2因子实验显示,假设排名稳定,但绝对评分高度依赖评估者身份,揭示单评判者大模型评估的局限性。
原文摘要 · Abstract (English)
Large language models have recently been explored for scientific hypothesis generation, but most prior work relies on unstructured literature and free-form textual claims. We present a pipeline for Earth observation that grounds hypothesis generation directly in the NASA Earth Observation Knowledge Graph. A heterogeneous graph neural network trained on historical co-usage relations ranks candidate dataset pairings, and a three-agent LLM pipeline filters, generates, and evaluates structured research hypotheses. Applied to 1,475 NASA datasets, the system produces 160 hypotheses spanning multiple Earth-science domains, including ecohydrology, glaciology, aerosol--cloud interactions, vegetation phenology, and stratospheric chemistry. Model-predicted novel dataset pairings are rated nearly as plausible as held-out real co-usages from the literature, indicating that the pipeline surfaces scientifically coherent yet unexplored combinations. A 2*2*2 factorial experiment across GPT-5.2 and Claude Sonnet 4.6 shows that hypothesis rankings remain stable, while absolute scores depend strongly on judge identity, highlighting limitations of single-judge LLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。