首个面向低空无人机图像指代分割的基准与模型,解决小物体密集场景下的语义理解难题。
RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation
- 分阶段注入语言信息,按类别和多尺度动态融合语义线索。
- 在13,871组真实无人机影像中验证,现有模型性能显著下降。
- 适合做无人机视觉-语言理解、复杂场景分割的研究者使用。
指代图像分割(RIS)旨在根据自然语言描述分割特定目标,在视觉-语言理解中具有重要意义。尽管遥感领域已有进展,低空无人机(LAD)场景下的RIS仍研究不足。现有数据集与方法多针对高空静态图像设计,难以应对LAD视角带来的多样视点与高密度物体问题。为此,我们提出RIS-LAD,首个专为LAD场景设计的细粒度RIS基准。该数据集包含从真实无人机视频中收集的13,871组图像-文本-掩码三元组,聚焦小物体、杂乱环境与多视角场景。其揭示了先前基准未涵盖的新挑战,如因小物体导致的类别漂移、同类密集物体间的对象漂移。为此,我们提出语义感知自适应推理网络(SAARN),不统一注入语言特征,而是将语义信息分解并路由至网络不同阶段:类别主导的语言增强(CDLE)在早期编码阶段对齐视觉特征与物体类别;自适应推理融合模块(ARFM)在多尺度下动态选择语义提示,提升复杂场景推理能力。实验表明,RIS-LAD对现有先进RIS算法构成显著挑战,且所提模型有效应对上述问题。数据集与代码将公开发布于:https://github.com/AHideoKuzeA/RIS-LAD/
原文摘要 · Abstract (English)
Referring Image Segmentation (RIS), which aims to segment specific objects based on natural language descriptions, plays an essential role in vision-language understanding. Despite its progress in remote sensing applications, RIS in Low-Altitude Drone (LAD) scenarios remains underexplored. Existing datasets and methods are typically designed for high-altitude and static-view imagery. They struggle to handle the unique characteristics of LAD views, such as diverse viewpoints and high object density. To fill this gap, we present RIS-LAD, the first fine-grained RIS benchmark tailored for LAD scenarios. This dataset comprises 13,871 carefully annotated image-text-mask triplets collected from realistic drone footage, with a focus on small, cluttered, and multi-viewpoint scenes. It highlights new challenges absent in previous benchmarks, such as category drift caused by tiny objects and object drift under crowded same-class objects. To tackle these issues, we propose the Semantic-Aware Adaptive Reasoning Network (SAARN). Rather than uniformly injecting all linguistic features, SAARN decomposes and routes semantic information to different stages of the network. Specifically, the Category-Dominated Linguistic Enhancement (CDLE) aligns visual features with object categories during early encoding, while the Adaptive Reasoning Fusion Module (ARFM) dynamically selects semantic cues across scales to improve reasoning in complex scenes. The experimental evaluation reveals that RIS-LAD presents substantial challenges to state-of-the-art RIS algorithms, and also demonstrates the effectiveness of our proposed model in addressing these challenges. The dataset and code will be publicly released soon at: https://github.com/AHideoKuzeA/RIS-LAD/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。