用可验证奖励训练模型,让AI从图示推断地质演化过程。
Geo-Strat-RL: Learning Geological Event Reasoning from Verifiable Tasks

- 设计可验证任务,用强化学习提升视觉语言模型推理能力。
- 在未见过的图示上,地质内容得分显著提高。
- 模型学到的推理能力可跨图示与地震数据通用。
为评估视觉-语言模型是否能推理地质历史,需构建已知过程历史的观测数据。地质推理不仅依赖视觉模式识别,还需理解难以直接观察或高度模糊的时间与结构关系。当真实事件历史不唯一或不可得时,如何训练模型生成符合观测证据和地质原理的合理重建仍是挑战。为此,我们提出Geo-Strat-RL,一个合成环境,可生成层序观测与紧凑的可见证据事件历史。该环境结合地质生成器与可执行验证器,对时间顺序、事件身份、沉积关系及结构关系进行评分。实验表明,使用可验证奖励的强化学习(RLVR)能提升视觉语言模型在新层序图上的地质内容得分。进一步在合成地震观测域中,将生成场景转化为声阻抗导出的振幅剖面,在受控配对渲染设置下,验证了从层序图训练所得的地质推理能力可迁移至地震表示,无需地震特定训练样本,支持了RLVR可传授跨观测格式的可复用地质推理概念的假设。
原文摘要 · Abstract (English)
To evaluate whether vision-language models can reason about geological histories, it is necessary to construct observations for which the underlying process history is known. Furthermore, reasoning over geological histories is not just a question of recognizing visual patterns, but also of understanding temporal and structural relationships that may be only indirectly visible or highly ambiguous. When ground-truth event histories are not uniquely identifiable or are unavailable, it remains an open challenge to teach models capable of visual reasoning to produce valid geological reconstructions that are consistent with both observed evidence and geological principles. We therefore investigate whether defining a verifiable geological reasoning task can improve geological event reconstruction across observation domains through reinforcement learning with verifiable rewards (RLVR). To this end, we present Geo-Strat-RL, a synthetic environment that generates stratigraphic observations and compact visible-evidence event histories. The environment combines a geological generator with an executable verifier that scores chronology, event identity, deposition, and structural relationships. We show that RLVR improves geological reconstruction in vision-language models (VLMs), increasing geological content scores on held out stratigraphic diagrams. We further evaluate the same held-out geological histories in a synthetic seismic observation domain by converting the generated scenes into acoustic-impedance-derived amplitude sections. In this controlled paired-renderer setting, we present evidence that geological reasoning learned from stratigraphic diagram-domain RLVR training transfers to synthetic seismic representations without seismic-specific training examples, supporting the hypothesis that RLVR can teach reusable geological reasoning concepts across related observation formats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。