用视频帧对数据集+高效微调,提升遥感时序变化检测能力
GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing
- 构建视频帧对数据集,捕捉地理变化动态
- 最佳模型达BERT分数0.864,精准描述土地利用变迁
- 适配遥感分析、环境监测等需要时序理解的场景
检测地理景观的时序变化对环境监测和城市规划至关重要。尽管遥感数据丰富,现有视觉语言模型(VLMs)难以有效捕捉时间动态。本文通过构建视频帧对标注数据集,追踪地理模式的演变。采用低秩适应(LoRA)、量化LoRA(QLoRA)及模型剪枝等微调技术,在Video-LLaVA和LLaVA-NeXT-Video等模型上显著提升处理遥感时序变化的能力。结果显示,最优模型在描述土地利用转换方面表现优异,达到BERT分数0.864和ROUGE-1分数0.576,验证了其高精度。
原文摘要 · Abstract (English)
Detecting temporal changes in geographical landscapes is critical for applications like environmental monitoring and urban planning. While remote sensing data is abundant, existing vision-language models (VLMs) often fail to capture temporal dynamics effectively. This paper addresses these limitations by introducing an annotated dataset of video frame pairs to track evolving geographical patterns over time. Using fine-tuning techniques like Low-Rank Adaptation (LoRA), quantized LoRA (QLoRA), and model pruning on models such as Video-LLaVA and LLaVA-NeXT-Video, we significantly enhance VLM performance in processing remote sensing temporal changes. Results show significant improvements, with the best performance achieving a BERT score of 0.864 and ROUGE-1 score of 0.576, demonstrating superior accuracy in describing land-use transformations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。