用语言提示缓解航拍图中光照视角等多重变化带来的检测难题
Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
- 通过视觉语义推理理解图像拍摄环境条件
- 设计关系学习损失函数,提升对视角尺度变化的鲁棒性
- 适合处理复杂场景下目标检测,尤其在多变航拍图像中
尽管计算机视觉取得进展,航拍图像中的目标检测仍面临诸多挑战,主要源于光照、视角等多种变化导致的图像场景差异大、物体外观剧烈改变。为此,本文提出一种名为LANGO的语言引导目标检测框架,通过语言引导学习缓解场景级与实例级变化的影响。首先,受人类感知环境因素(如天气)理解语义方式启发,设计视觉语义推理模块,解析图像拍摄条件以理解场景语义;其次,提出关系学习损失函数,利用语言表征中对视角、尺度变化具有鲁棒性的特性,学习类别间的关系,以应对实例级变化。大量实验表明,该方法显著提升检测性能。
原文摘要 · Abstract (English)
Despite recent advancements in computer vision research, object detection in aerial images still suffers from several challenges. One primary challenge to be mitigated is the presence of multiple types of variation in aerial images, for example, illumination and viewpoint changes. These variations result in highly diverse image scenes and drastic alterations in object appearance, so that it becomes more complicated to localize objects from the whole image scene and recognize their categories. To address this problem, in this paper, we introduce a novel object detection framework in aerial images, named LANGuage-guided Object detection (LANGO). Upon the proposed language-guided learning, the proposed framework is designed to alleviate the impacts from both scene and instance-level variations. First, we are motivated by the way humans understand the semantics of scenes while perceiving environmental factors in the scenes (e.g., weather). Therefore, we design a visual semantic reasoner that comprehends visual semantics of image scenes by interpreting conditions where the given images were captured. Second, we devise a training objective, named relation learning loss, to deal with instance-level variations, such as viewpoint angle and scale changes. This training objective aims to learn relations in language representations of object categories, with the help of the robust characteristics against such variations. Through extensive experiments, we demonstrate the effectiveness of the proposed method, and our method obtains noticeable detection performance improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。