arXiv:2411.09180cs.CVcs.AI2024-11中稿 · paper

用可学习提示提升无人机图像目标检测的泛化能力

LEAP:D -- A Novel Prompt-based Approach for Domain-Generalized Aerial Object Detection

  • 引入可学习提示替代人工提示,减少领域偏差
  • 一步训练法同步优化提示与模型,效率更高
  • 显著增强模型在多变环境下的鲁棒性,适合实际航拍场景

无人机拍摄图像因飞行高度、角度和天气等条件变化,导致目标外观和形状差异显著,影响目标检测性能。为应对这一挑战,本文提出一种基于可学习提示的视觉-语言新方法。该方法取代传统手动设计提示,降低特定领域知识干扰,提升模型泛化能力。同时采用一步式训练策略,在模型训练过程中同步更新可学习提示,提升训练效率且不牺牲性能。实验表明,该方法有效增强了模型在多样化环境中的鲁棒性与适应性,推动了领域泛化目标检测的发展。

原文摘要 · Abstract (English)

Drone-captured images present significant challenges in object detection due to varying shooting conditions, which can alter object appearance and shape. Factors such as drone altitude, angle, and weather cause these variations, influencing the performance of object detection algorithms. To tackle these challenges, we introduce an innovative vision-language approach using learnable prompts. This shift from conventional manual prompts aims to reduce domain-specific knowledge interference, ultimately improving object detection capabilities. Furthermore, we streamline the training process with a one-step approach, updating the learnable prompt concurrently with model training, enhancing efficiency without compromising performance. Our study contributes to domain-generalized object detection by leveraging learnable prompts and optimizing training processes. This enhances model robustness and adaptability across diverse environments, leading to more effective aerial object detection.

目标检测无人机图像提示学习域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。