首个面向全球梯田地物提取的多模态数据集,融合影像、文本与高程信息。
GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality

- 构建包含影像、文本和高程数据的多模态梯田数据集
- 引入文本与高程信息后,分割准确率提升12.3%且边界更连贯
- 适合遥感农业监测、地形感知分割研究者使用
农田地物提取在基于遥感的农业监测中至关重要,支持地块调查、精准管理与生态评估。然而现有公开基准多聚焦于规则平坦农田,而山地梯田具有阶梯状地形、显著高程变化、边界不规则及跨区域异质性强等特点,使地物提取更具挑战性,需融合视觉识别、语义区分与地形感知几何理解。尽管近期研究推进了视觉地物基准与图文农田理解,但缺乏在统一图像-文本-数字高程模型(DEM)设置下的梯田提取基准。为此,我们提出GTPBD-MM,首个面向全球梯田地物提取的多模态基准。基于GTPBD,GTPBD-MM整合高分辨率光学影像、结构化文本描述与DEM数据,支持在仅图像、图像+文本、图像+文本+DEM三种设置下的系统评估。我们进一步提出高程与文本引导的梯田分割网络ETTerra作为多模态基线。大量实验表明,文本语义与地形几何提供超越视觉外观的互补线索,在复杂梯田场景中实现更准确、连贯、结构一致的分割结果。
原文摘要 · Abstract (English)
Agricultural parcel extraction plays an important role in remote sensing-based agricultural monitoring, supporting parcel surveying, precision management, and ecological assessment. However, existing public benchmarks mainly focus on regular and relatively flat farmland scenes. In contrast, terraced parcels in mountainous regions exhibit stepped terrain, pronounced elevation variation, irregular boundaries, and strong cross-regional heterogeneity, making parcel extraction a more challenging problem that jointly requires visual recognition, semantic discrimination, and terrain-aware geometric understanding. Although recent studies have advanced visual parcel benchmarks and image-text farmland understanding, a unified benchmark for complex terraced parcel extraction under aligned image-text-DEM settings remains absent. To fill this gap, we present GTPBD-MM, the first multimodal benchmark for global terraced parcel extraction. Built upon GTPBD, GTPBD-MM integrates high-resolution optical imagery, structured text descriptions, and DEM data, and supports systematic evaluation under Image-only, Image+Text, and Image+Text+DEM settings. We further propose Elevation and Text guided Terraced parcel network (ETTerra), a multimodal baseline for terraced parcel delineation. Extensive experiments demonstrate that textual semantics and terrain geometry provide complementary cues beyond visual appearance alone, yielding more accurate, coherent, and structurally consistent delineation results in complex terraced scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。