arXiv:2502.16469cs.CVcs.AI2025-02IJCV被引 2

用丰富文本信息提升跨域少样本目标检测性能

Cross-domain Few-shot Object Detection with Multi-modal Textual Enrichment

  • 基于元学习框架,融合多模态文本增强视觉特征对齐
  • 在多个跨域基准上超越现有少样本检测方法
  • 适合关注跨域适应与文本增强的计算机视觉研究者

跨模态特征提取与融合的进步显著提升了少样本学习性能。然而,当前多模态目标检测(MM-OD)方法在面临显著领域差异时仍会出现明显性能下降。我们提出,引入丰富的文本信息可帮助模型建立更稳健的视觉实例与语言描述之间的知识关联,从而缓解领域漂移问题。针对跨域多模态少样本目标检测(CDMM-FSOD)问题,我们提出一种基于元学习的框架,利用丰富文本语义作为辅助模态实现有效领域自适应。新架构包含两个核心组件:(i) 多模态特征聚合模块,对齐视觉与语言特征嵌入,确保模态间一致融合;(ii) 丰富文本语义校正模块,通过双向文本特征生成优化多模态特征对齐,提升语言理解及其在目标检测中的应用效果。我们在常见跨域目标检测基准上评估该方法,结果表明其显著优于现有少样本目标检测方法。

原文摘要 · Abstract (English)

Advancements in cross-modal feature extraction and integration have significantly enhanced performance in few-shot learning tasks. However, current multi-modal object detection (MM-OD) methods often experience notable performance degradation when encountering substantial domain shifts. We propose that incorporating rich textual information can enable the model to establish a more robust knowledge relationship between visual instances and their corresponding language descriptions, thereby mitigating the challenges of domain shift. Specifically, we focus on the problem of Cross-Domain Multi-Modal Few-Shot Object Detection (CDMM-FSOD) and introduce a meta-learning-based framework designed to leverage rich textual semantics as an auxiliary modality to achieve effective domain adaptation. Our new architecture incorporates two key components: (i) A multi-modal feature aggregation module, which aligns visual and linguistic feature embeddings to ensure cohesive integration across modalities. (ii) A rich text semantic rectification module, which employs bidirectional text feature generation to refine multi-modal feature alignment, thereby enhancing understanding of language and its application in object detection. We evaluate the proposed method on common cross-domain object detection benchmarks and demonstrate that it significantly surpasses existing few-shot object detection approaches.

少样本检测跨域适应多模态融合文本增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。