用大模型让卫星图自动分析森林变化,支持自然语言提问。
Forest-Chat: Adapting Vision-Language Agents for Interactive Forest Change Analysis
- 构建多层级视觉语言框架,结合大模型实现零样本变化检测与描述。
- 在自建数据集上达到67.10% mIoU和40.17% BLEU-4,零样本仍表现稳健。
- 适合需要可解释森林变化分析的科研与环保工作者使用。
高分辨率卫星影像与深度学习的发展为森林监测带来新机遇。本文提出Forest-Chat,一种基于大语言模型(LLM)的交互式森林变化分析代理,支持自然语言查询,涵盖变化检测、图像描述、物体计数、毁林特征分析与变化推理等任务。系统采用多层级变化解释(MCI)视觉语言骨干网络,结合零样本变化检测(AnyChange)与多模态大模型驱动的零样本描述生成与优化。为支持森林场景评估,构建了包含双时相影像、像素级变化掩码及语义变化描述的Forest-Change数据集,由人工标注与规则方法生成。Forest-Chat在该数据集上获得67.10% mIoU和40.17% BLEU-4,在LEVIR-MCI-Trees子集上达88.13%与34.41%;零样本条件下分别取得60.15%与34.00%(Forest-Change)、47.32%与18.23%(LEVIR-MCI-Trees)。实验表明,描述优化能有效注入地理领域知识,但标签迁移能力有限于JL1-CD-Trees。结果证明,交互式大模型系统可实现可访问且可解释的森林变化分析。
原文摘要 · Abstract (English)
The increasing availability of high-resolution satellite imagery, together with advances in deep learning, creates new opportunities for forest monitoring workflows. Two central challenges in this domain are pixel-level change detection and semantic change interpretation, particularly for complex forest dynamics. While large language models (LLMs) are increasingly adopted for data exploration, their integration with vision-language models (VLMs) for remote sensing image change interpretation (RSICI) remains underexplored, especially beyond urban environments. This paper introduces Forest-Chat, an LLM-driven agent for forest change analysis, enabling natural language querying across multiple RSICI tasks, including change detection and captioning, object counting, deforestation characterisation, and change reasoning. Forest-Chat builds upon a multi-level change interpretation (MCI) vision-language backbone with LLM-based orchestration, incorporating zero-shot change detection via AnyChange and multimodal LLM-based zero-shot change captioning and refinement. To support adaptation and evaluation in forest environments, we introduce the Forest-Change dataset, comprising bi-temporal satellite imagery, pixel-level change masks, and semantic change captions via human annotation and rule-based methods. Forest-Chat achieves mIoU and BLEU-4 scores of 67.10% and 40.17% on Forest-Change, and 88.13% and 34.41% on LEVIR-MCI-Trees, a tree-focused subset of LEVIR-MCI. In a zero-shot capacity, it achieves 60.15% and 34.00% on Forest-Change, and 47.32% and 18.23% on LEVIR-MCI-Trees. Further experiments demonstrate the value of caption refinement for injecting geographic domain knowledge into supervised captions, and the system's limited label domain transfer onto JL1-CD-Trees. These findings demonstrate that interactive, LLM-driven systems can support accessible and interpretable forest change analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。