用多模态数据与领域规则增强智能体,实现滑坡自动识别与分析。
LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis

- 构建滑坡专用多模态数据集,含像素级标注与高质量文本描述。
- 模型在细粒度分类上准确率提升32.87%,语义描述质量显著改善。
- 适合地质灾害监测、遥感分析等领域的研究与应用者使用。
智能滑坡灾害解析对防灾至关重要,但现有方法难以同时提取视觉特征与高阶地质语义,通用视觉语言模型在复杂地质场景中存在感知局限与领域幻觉问题。为此,我们提出一种指令驱动的智能体框架,包含三个部分:首先,通过多视觉语言模型交叉验证与交互标注,构建包含七种亚型标签、高分辨率图像、像素级掩码及高质量文本描述的多模态细粒度数据集LandslideBench;其次,基于该数据集,使用LoRA微调得到面向滑坡的视觉语言模型LandslideVLM,增强地质语义理解能力;最后,设计以LandslideVLM为认知核心的领域规则增强型智能体LandslideAgent,采用双规则控制器,结合结构化报告元数据约束与交叉验证识别约束,调控自动化工具调用。实验表明,LandslideBench为五种主流模型提供了有效的基准,在细粒度分类与语义分割任务上表现优异。LandslideVLM在滑坡判别、细粒度分类与语义描述质量上分别提升10.96%、32.87%和15.91%。LandslideAgent进一步实现多源空间数据的自主推理,完成滑坡识别与分析全流程智能化。
原文摘要 · Abstract (English)
Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and high-level geoscientific semantics, while general-purpose vision-language models (VLMs) suffer from perceptual limitations and domain hallucinations in complex geological scenarios. To address these challenges, we propose an instruction-driven agentic framework comprising three components. First, LandslideBench, a multimodal fine-grained dataset with seven subtype labels, high-resolution imagery, pixel-level masks, and high-quality textual descriptions, is constructed via multi-VLM cross-validation and interactive annotation. Then, LandslideVLM, a landslide-oriented VLM, is fine-tuned via LoRA on LandslideBench to enhance geological semantic understanding. Finally, LandslideAgent, a domain rule-enhanced agent taking LandslideVLM as its cognitive backbone, employs a dual-rule controller incorporating structured report metadata constraints and cross-validation identification constraints to regulate automated tool invocation. Experiments demonstrate that LandslideBench provides effective baselines across five mainstream models on fine-grained classification and semantic segmentation. LandslideVLM achieves accuracy improvements of 10.96%, 32.87%, and 15.91% on landslide discrimination, fine-grained classification, and semantic description quality, respectively. LandslideAgent further enables autonomous multi-source spatial data inference, realizing full-process intelligence for landslide identification and analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。