首个面向农田时空分割的图文数据集,支持语言驱动学习。
A large-scale image-text dataset benchmark for farmland segmentation
- 提出半自动标注方法,高效生成高质量图文对。
- 覆盖中国8大农业区、四季变化,包含丰富时空特征描述。
- 适合遥感、农业AI研究者使用,推动农田动态建模研究。
传统深度学习依赖标注数据,在表达农田要素与环境的空间关系及动态演变方面存在局限,难以有效建模农田的时空异质性。语言作为结构化知识载体,可显式表达农田的形状、分布及周边环境等时空特征,因此语言驱动学习能有效缓解这一挑战。然而,当前遥感农田领域尚无全面的基准数据集支持该方向研究。为此,本文引入基于语言的农田描述,构建了首个细粒度图像-文本数据集FarmSeg-VL,用于时空农田分割。首先,提出一种半自动标注方法,可精准为每张图像分配文本描述,兼顾数据质量与语义丰富性,同时提升构建效率。其次,FarmSeg-VL具备显著的时空特性:时间维度涵盖四季;空间维度覆盖中国八大典型农业区域;文本内容涵盖农田本体属性、物候特征、空间分布、地形地貌及周边环境分布等。最后,对视觉语言模型(VLMs)及仅依赖标签的深度学习模型在该数据集上的表现进行了分析,验证其作为农田分割标准基准的潜力。
原文摘要 · Abstract (English)
The traditional deep learning paradigm that solely relies on labeled data has limitations in representing the spatial relationships between farmland elements and the surrounding environment.It struggles to effectively model the dynamic temporal evolution and spatial heterogeneity of farmland. Language,as a structured knowledge carrier,can explicitly express the spatiotemporal characteristics of farmland, such as its shape, distribution,and surrounding environmental information.Therefore,a language-driven learning paradigm can effectively alleviate the challenges posed by the spatiotemporal heterogeneity of farmland.However,in the field of remote sensing imagery of farmland,there is currently no comprehensive benchmark dataset to support this research direction.To fill this gap,we introduced language based descriptions of farmland and developed FarmSeg-VL dataset,the first fine-grained image-text dataset designed for spatiotemporal farmland segmentation.Firstly, this article proposed a semi-automatic annotation method that can accurately assign caption to each image, ensuring high data quality and semantic richness while improving the efficiency of dataset construction.Secondly,the FarmSeg-VL exhibits significant spatiotemporal characteristics.In terms of the temporal dimension,it covers all four seasons.In terms of the spatial dimension,it covers eight typical agricultural regions across China.In addition, in terms of captions,FarmSeg-VL covers rich spatiotemporal characteristics of farmland,including its inherent properties,phenological characteristics, spatial distribution,topographic and geomorphic features,and the distribution of surrounding environments.Finally,we present a performance analysis of VLMs and the deep learning models that rely solely on labels trained on the FarmSeg-VL,demonstrating its potential as a standard benchmark for farmland segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。