用提示驱动生成多样样本,提升少样本跨域目标检测性能
Prompt-Driven Simulation with Feature Perturbation for Cross-Domain Few-Shot Object Detection

- 通过大视觉模型提示控制前景背景变化,生成语义一致的多样化样本
- 引入特征扰动机制,使模型在噪声下保持稳定预测,减少对领域线索依赖
- 适合解决标注少、领域差异大的目标检测场景
数据增强通过模拟多样视觉变化扩展源域分布并诱导合成领域偏移,是缓解跨域少样本目标检测中严重领域差异和标注数据不足的有效方法。现有方法依赖传统增强技术(如颜色抖动、拼接、背景中心适应等),难以建模复杂领域偏移,常导致性能不佳。本文提出PSP-FSOD框架,将提示驱动的领域模拟与特征扰动正则化结合,提升跨域少样本目标检测的泛化能力。为实现可控领域合成,设计基于大视觉语言模型提示的策略,联合建模前景与背景变化,生成语义一致且领域多样的训练样本;同时采用感知基础的生成方案,指导物体位置布局,缓解语义-空间错位问题,改善前景适应性。为保障训练稳定性与鲁棒性,进一步引入高斯噪声注入的多尺度中间特征扰动机制,并进行分布校正,促使模型在扰动下保持一致预测,降低对领域特有线索的依赖。大量实验表明,PSP-FSOD能生成高质量的领域多样化监督信号,学习到领域不变表示,在多个跨域少样本目标检测基准上持续提升性能。
原文摘要 · Abstract (English)
Data augmentation, which simulates diverse visual variations to expand the source distribution and induce synthetic domain shifts, is a simple yet effective strategy for mitigating severe domain shifts and limited labeled target data in cross-domain few-shot object detection (CD-FSOD). Existing approaches rely on conventional data augmentation, such as Color-Jitter, Mosaic, and background-centric adaptation (e.g., Domain-RAG), which are limited in modeling complex domain shifts and often lead to suboptimal performance. In this paper, we propose PSP-FSOD, a principled framework that integrates prompt-driven domain simulation with feature perturbation regularization to improve generalization in CD-FSOD. To enable controllable domain synthesis, we design a prompt-driven strategy that leverages the visual grounding capability of large VLMs to jointly model foreground and background variations, generating semantically consistent yet domain-diverse training samples. Moreover, we adopt a grounding-aware generation scheme that guides object placement and alleviates semantic-spatial misalignment, thereby improving foreground adaptation. To ensure training stability and robustness, we further introduce a noise-induced feature perturbation mechanism that injects Gaussian noise into multi-scale intermediate features with distribution correction, encouraging consistent predictions under perturbations and reducing reliance on domain-specific cues. Extensive experiments demonstrate that PSP-FSOD produces high-quality domain-diverse supervision and learns domain-invariant representations, consistently improving performance across CD-FSOD benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。