用多模态模型自动解读月球地质,让机器推理有据可查。
Verifiably grounded machine interpretation of lunar geology
- 通过融合地形、光谱与地质图,模型从局部数据推断地层结构。
- 仅靠视觉信息推算年龄会陷入记忆偏差,准确率低。
- 引入开放书检索机制后,能引用文献年代,实现可信推理。
行星地质学依赖历史性的解释推理,从多样观测中复原过去事件。本文研究多模态视觉-语言模型在多大程度上可自动化这一流程。聚焦月球玄武岩熔岩平原的地层结构,我们训练模型直接从配准的地形、光谱和地质图生成可验证的地质解释。结果表明,该系统能有效平衡已知地质先验与局部视觉证据,准确描述地层与地形特征;但仅依赖视觉信息推算年龄时,模型会偏向记忆中的先验,导致误差。引入开放书式检索机制后,模型可准确引用已有文献中的年代数据。研究揭示了自动化地质推断所需架构:局部现场证据需由视觉数据解释,而定量历史背景则须从科学文献中检索获取。
原文摘要 · Abstract (English)
Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. Here, we investigate how far this interpretive workflow can be automated by a multimodal vision-language model. Focusing on the stratigraphy of lunar basaltic mare volcanism, we train a model to generate verifiably grounded geologic interpretations directly from co-registered topographic, spectral, and geologic maps. We demonstrate that while the system successfully balances established geological priors with local visual evidence to accurately describe stratigraphy and terrain, numeric age dating derived solely from vision defaults to memorized priors. Integrating an open-book retrieval mechanism resolves this, enabling the model to faithfully cite published chronologies. Our findings delineate the necessary architecture for automated geologic inference: site evidence must be visually interpreted from local data, while quantitative historical context must be retrieved from the scientific record.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。