arXiv:2503.16106cs.CV2025-03CVPR被引 8

解决少样本下开放集域泛化难题,提升CLIP模型在数据稀缺时的鲁棒性。

OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP

  • 设计跨注意力模块融合视觉与语义提示,增强跨模态适应能力。
  • 通过合成伪开放样本训练专用提示,显著提升对未知类别的检测精度。
  • 适用于少样本场景下的开放集识别任务,尤其适合资源受限的部署环境。

我们提出低样本开放集域泛化(LSOSDG)新范式,统一少样本学习与开放集域泛化。尽管基于CLIP的提示方法推动了域泛化发展,但在低数据场景(如1样本)下表现不佳,且难以精准识别与训练类别具有细粒度语义关联的开放集样本。为此,我们提出OSLOPROMPT框架,包含两项核心创新:第一,引入无领域依赖的提示学习机制,通过新颖的交叉注意力模块整合可调领域特异性线索与视觉引导语义属性,并辅以可学习的通用视觉提示,提升跨模态适应性;第二,为改善推理阶段异常值拒绝能力,将未知样本分类为“未知”,并利用现成基础模型通过定向查询策略生成保持与已知类别细粒度关系的伪开放样本,系统性训练专用提示,增强特征学习能力,使模型更有效检测不同粒度的开放样本。在五个基准上的大量实验表明,OSLOPROMPT在LSOSDG上达到新最优性能,显著优于现有方法。

原文摘要 · Abstract (English)

We introduce Low-Shot Open-Set Domain Generalization (LSOSDG), a novel paradigm unifying low-shot learning with open-set domain generalization (ODG). While prompt-based methods using models like CLIP have advanced DG, they falter in low-data regimes (e.g., 1-shot) and lack precision in detecting open-set samples with fine-grained semantics related to training classes. To address these challenges, we propose OSLOPROMPT, an advanced prompt-learning framework for CLIP with two core innovations. First, to manage limited supervision across source domains and improve DG, we introduce a domain-agnostic prompt-learning mechanism that integrates adaptable domain-specific cues and visually guided semantic attributes through a novel cross-attention module, besides being supported by learnable domain- and class-generic visual prompts to enhance cross-modal adaptability. Second, to improve outlier rejection during inference, we classify unfamiliar samples as "unknown" and train specialized prompts with systematically synthesized pseudo-open samples that maintain fine-grained relationships to known classes, generated through a targeted query strategy with off-the-shelf foundation models. This strategy enhances feature learning, enabling our model to detect open samples with varied granularity more effectively. Extensive evaluations across five benchmarks demonstrate that OSLOPROMPT establishes a new state-of-the-art in LSOSDG, significantly outperforming existing methods.

CLIP少样本学习开放集识别提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。