用视觉提示提升遥感图像目标检测的开放词汇能力
Beyond Open Vocabulary: Multimodal Prompting for Object Detection in Remote Sensing Images
- 引入视觉提示编码器,实现无需文本的类别指定
- 多模态融合使文本与视觉信息协同,提升检测稳定性
- 在语义模糊和分布偏移下仍保持可靠性能,适合复杂遥感场景
遥感图像中的开放词汇目标检测通常依赖纯文本提示来指定目标类别,隐含假设推理时的类别查询可通过预训练诱导的图文对齐可靠定位。但在实际遥感场景中,因任务与应用特有的类别语义,该假设常失效,导致开放词汇设置下类别指定不稳定。为此,本文提出RS-MPOD,一种基于多模态提示的开放词汇检测框架,通过引入实例级视觉提示、文本提示及其多模态融合,突破纯文本提示限制。该框架包含视觉提示编码器,从样本实例中提取基于外观的类别线索,实现无文本类别指定;以及多模态融合模块,在双模态可用时整合信息。在标准、跨数据集及细粒度遥感基准上的实验表明,视觉提示在语义模糊和分布偏移下更可靠,而多模态提示在文本语义对齐良好时仍具竞争力。
原文摘要 · Abstract (English)
Open-vocabulary object detection in remote sensing commonly relies on text-only prompting to specify target categories, implicitly assuming that inference-time category queries can be reliably grounded through pretraining-induced text-visual alignment. In practice, this assumption often breaks down in remote sensing scenarios due to task- and application-specific category semantics, resulting in unstable category specification under open-vocabulary settings. To address this limitation, we propose RS-MPOD, a multimodal open-vocabulary detection framework that reformulates category specification beyond text-only prompting by incorporating instance-grounded visual prompts, textual prompts, and their multimodal integration. RS-MPOD introduces a visual prompt encoder to extract appearance-based category cues from exemplar instances, enabling text-free category specification, and a multimodal fusion module to integrate visual and textual information when both modalities are available. Extensive experiments on standard, cross-dataset, and fine-grained remote sensing benchmarks show that visual prompting yields more reliable category specification under semantic ambiguity and distribution shifts, while multimodal prompting provides a flexible alternative that remains competitive when textual semantics are well aligned.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。