针对无人机动态视角下的场景变化,生成更准确的语义描述。
Hierarchical Dual-Change Collaborative Learning for UAV Scene Change Captioning
- 设计动态自适应布局注意力机制,统一建模重叠与非重叠区域特征。
- 提出分层跨模态方向一致性校准,提升对相机移动方向的敏感性。
- 构建首个无人机场景变化描述数据集,推动该任务研究。
本文提出一项新任务——无人机场景变化描述(UAV-SCC),旨在为移动视角下动态航拍图像对中的语义变化生成自然语言描述。不同于传统固定视角图像对的变化描述,UAV-SCC需处理因相机移动导致的时空双重变化,且两帧图像仅部分重叠。为此,我们提出分层双变化协同学习(HDC-CL)方法,设计新型动态自适应布局变换器(DALT),在统一编码层中自适应建模图像对的多样空间布局,有效融合重叠与非重叠区域的相关特征;同时提出分层跨模态方向一致性校准(HCM-OCC)方法,增强模型对视角偏移方向的感知能力,提升变化描述精度。为推进该任务研究,我们构建了首个专用基准数据集UAV-SCC dataset。大量实验表明,所提方法在该任务上达到最先进性能。数据集与代码将在论文录用后公开。
原文摘要 · Abstract (English)
This paper proposes a novel task for UAV scene understanding - UAV Scene Change Captioning (UAV-SCC) - which aims to generate natural language descriptions of semantic changes in dynamic aerial imagery captured from a movable viewpoint. Unlike traditional change captioning that mainly describes differences between image pairs captured from a fixed camera viewpoint over time, UAV scene change captioning focuses on image-pair differences resulting from both temporal and spatial scene variations dynamically captured by a moving camera. The key challenge lies in understanding viewpoint-induced scene changes from UAV image pairs that share only partially overlapping scene content due to viewpoint shifts caused by camera rotation, while effectively exploiting the relative orientation between the two images. To this end, we propose a Hierarchical Dual-Change Collaborative Learning (HDC-CL) method for UAV scene change captioning. In particular, a novel transformer, \emph{i.e.} Dynamic Adaptive Layout Transformer (DALT) is designed to adaptively model diverse spatial layouts of the image pair, where the interrelated features derived from the overlapping and non-overlapping regions are learned within the flexible and unified encoding layer. Furthermore, we propose a Hierarchical Cross-modal Orientation Consistency Calibration (HCM-OCC) method to enhance the model's sensitivity to viewpoint shift directions, enabling more accurate change captioning. To facilitate in-depth research on this task, we construct a new benchmark dataset, named UAV-SCC dataset, for UAV scene change captioning. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on this task. The dataset and code will be publicly released upon acceptance of this paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。