arXiv:2508.09560cs.CVcs.RO2025-08NeurIPS被引 11

通过图文融合提升无人机在各种天气下的定位精度。

WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization

  • 用预训练大模型生成多天气文本描述,实现无训练天气推理。
  • 通过动态门控机制分离场景与天气特征,提升表示解耦能力。
  • 适合需要强泛化能力的无人机导航系统研究者使用。

无人机视觉地理定位在雨雪、雾霾等天气下性能显著下降,现有方法存在两大局限:一是依赖有限天气类别,泛化能力差;二是通过伪天气类别对纠缠的场景-天气特征解耦效果不佳。本文提出WeatherPrompt,一种多模态学习范式,通过融合图像嵌入与文本上下文建立天气不变表示。首先,引入无需训练的天气推理机制,利用现成的大规模多模态模型通过类人推理生成多天气文本描述,提升对未见或复杂天气的可扩展性,并能反映不同天气强度。其次,提出基于文本嵌入驱动的动态门控机制,自适应重加权并融合跨模态视觉特征,更好解耦场景与天气信息。框架还通过跨模态目标优化,包括图像-文本对比学习和匹配,使同一场景在不同天气下的表示更接近。大量实验表明,该方法在多种天气条件下表现优异,相比现有顶尖无人机定位方法,夜间条件下Recall@1提升13.37%,雾天和雪天提升18.69%。

原文摘要 · Abstract (English)

Visual geo-localization for drones faces critical degradation under weather perturbations, \eg, rain and fog, where existing methods struggle with two inherent limitations: 1) Heavy reliance on limited weather categories that constrain generalization, and 2) Suboptimal disentanglement of entangled scene-weather features through pseudo weather categories. We present WeatherPrompt, a multi-modality learning paradigm that establishes weather-invariant representations through fusing the image embedding with the text context. Our framework introduces two key contributions: First, a Training-free Weather Reasoning mechanism that employs off-the-shelf large multi-modality models to synthesize multi-weather textual descriptions through human-like reasoning. It improves the scalability to unseen or complex weather, and could reflect different weather strength. Second, to better disentangle the scene and weather feature, we propose a multi-modality framework with the dynamic gating mechanism driven by the text embedding to adaptively reweight and fuse visual features across modalities. The framework is further optimized by the cross-modal objectives, including image-text contrastive learning and image-text matching, which maps the same scene with different weather conditions closer in the respresentation space. Extensive experiments validate that, under diverse weather conditions, our method achieves competitive recall rates compared to state-of-the-art drone geo-localization methods. Notably, it improves Recall@1 by +13.37\% under night conditions and by 18.69\% under fog and snow conditions.

无人机定位多模态学习天气鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。