用高效微调让小模型具备地质推理专家水平。
Geo-Expert: Towards Expert-Level Geological Reasoning via Parameter-Efficient Fine-Tuning
- 用自研指令数据与低秩适配微调小模型,提升地质推断能力。
- 8B模型在专业评测中超越70B通用大模型和GPT-4o。
- 适合地质科研、智能勘探等需要高精度推理的场景。
通用大语言模型在地质学应用中常因对地下结构和深时演化推理产生幻觉而受限,当前AI在地球科学领域多集中于地表遥感与GIS。为填补这一空白,我们提出Geo-Expert,一个基于自定义高质量指令数据集(通过自研指令合成流程构建)微调的参数高效地质大模型系列。通过低秩适配(LoRA)方法,对Qwen3-8B、Qwen3-32B和Gemma-3-27B三个基础模型进行微调,并研究模型规模与架构的影响。在新提出的领域专用基准Geo-Eval上的广泛评估显示,经过领域对齐的8B模型在专业地质推理任务上超越开源70B通用模型及专有GPT-4o,而32B版本接近前沿推理模型表现。优化后的8B模型还展现出优异的成本-性能比,适用于实际部署。本工作提供可复现的科学大模型民主化路径,并建立地质人工智能的基准线。
原文摘要 · Abstract (English)
While general-purpose Large Language Models (LLMs) applied to Geology often hallucinate when reasoning about subsurface structures and deep-time evolution, current AI in Earth sciences predominantly targets surface remote sensing and GIS. To bridge this gap, we introduce Geo-Expert, a family of parameter-efficient geological LLMs fine-tuned on a custom-curated, high-quality instruction dataset processed using our custom instruction synthesis pipeline. We investigate the impact of model scaling and architecture by fine-tuning three base models: Qwen3-8B, Qwen3-32B, and Gemma-3-27B, with Low-Rank Adaptation (LoRA) method. Our extensive evaluation on a novel domain-specific benchmark, Geo-Eval, reveals that a domain-aligned 8B model can outperform open-weight 70B generalists and proprietary GPT-4o on specialized geological reasoning, while a 32B variant approaches frontier reasoning models. The optimized 8B model further offers a competitive cost-performance ratio for deployment. This work provides a reproducible recipe for democratizing scientific LLMs and establishes a baseline for geological artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。