提出新模型与数据集,让无人机图像自动描述城市施工变化。
UAV as Urban Construction Change Monitor: A New Benchmark and Change Captioning Model

- 用可学习原型库显式建模变化语义,统一变化检测与描述生成。
- 在9000对高清航拍图上实现更精准的语义变化描述,优于现有方法。
- 适合做城市监测、遥感智能分析的研究者和开发者使用。
遥感图像变化描述(RSICC)旨在从双时相影像中生成具有空间定位的自然语言描述,推动从二值变化掩码迈向语义级理解。然而现有方法依赖隐式特征差分,未显式建模结构化变化语义,且难以平衡变化检测与文本生成的表示需求。同时,现有基准对高分辨率城市施工场景覆盖有限。为此,我们提出PTNet——一种原型引导的任务自适应框架,通过可学习原型库指导跨时相交互,利用多头门控解耦任务特定表示,并将检测所得空间先验注入生成过程,实现语义一致性与细粒度空间敏感性的兼顾。此外,构建了基于无人机的大规模基准UCCD,包含9000对高分辨率图像对和45,000条标注语句,用于城市施工监测。在UCCD与WHU-CDC上的实验表明,PTNet持续优于现有方法。数据集与代码已公开于https://github.com/G124556/ptnet。
原文摘要 · Abstract (English)
Remote Sensing Image Change Captioning (RSICC) aims to generate spatially grounded natural language descriptions of scene evolution from bi-temporal imagery, moving beyond binary change masks toward semantic-level understanding. However, existing methods rely on implicit feature differencing without explicitly modeling structured change semantics, and struggle to reconcile the conflicting representation demands of change detection and caption generation. In addition, current benchmarks provide limited coverage of high-resolution urban construction scenarios. To address these challenges, we propose PTNet, a prototype-guided task-adaptive framework for joint change captioning and detection. PTNet explicitly models structured change semantics through a learnable prototype bank that guides cross-temporal interaction, disentangles task-specific representations via multi-head gating, and injects detection-derived spatial priors into caption generation, enabling coherent semantic correspondence while preserving fine-grained spatial sensitivity. Furthermore, we construct UCCD, a large-scale UAV-based benchmark comprising 9,000 high-resolution image pairs and 45,000 annotated sentences for urban construction monitoring. Extensive experiments on UCCD and WHU-CDC demonstrate that PTNet consistently outperforms existing methods. The dataset and source code are publicly available at https://github.com/G124556/ptnet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。