arXiv:2511.10352cs.CV2025-11

用傅里叶与方向分布增强目标检测的域泛化能力

FOUND: Fourier-based von Mises Distribution for Robust Single Domain Generalization in Object Detection

  • 用vMF分布建模特征方向,捕捉语义不变结构
  • 在频域扰动幅度相位,模拟域偏移提升鲁棒性
  • 适合需要强泛化的自动驾驶目标检测场景

单域泛化(SDG)目标检测旨在仅用单一源域训练模型,使其有效适应未见目标域。尽管基于CLIP的语义增强方法已有进展,但常忽视特征分布与频域特性对鲁棒性的重要影响。本文提出新框架,将von Mises-Fisher(vMF)分布与傅里叶变换融入CLIP引导流程。通过vMF建模对象表征的方向特征,更好捕捉嵌入空间中的域不变语义结构;同时引入基于傅里叶的增强策略,扰动频域中的幅度与相位成分,模拟域偏移以提升特征鲁棒性。该方法在保持CLIP语义对齐优势的同时,增强了特征多样性与跨域结构一致性。在多样天气驾驶基准上的大量实验表明,本方法优于现有最先进方法。

原文摘要 · Abstract (English)

Single Domain Generalization (SDG) for object detection aims to train a model on a single source domain that can generalize effectively to unseen target domains. While recent methods like CLIP-based semantic augmentation have shown promise, they often overlook the underlying structure of feature distributions and frequency-domain characteristics that are critical for robustness. In this paper, we propose a novel framework that enhances SDG object detection by integrating the von Mises-Fisher (vMF) distribution and Fourier transformation into a CLIP-guided pipeline. Specifically, we model the directional features of object representations using vMF to better capture domain-invariant semantic structures in the embedding space. Additionally, we introduce a Fourier-based augmentation strategy that perturbs amplitude and phase components to simulate domain shifts in the frequency domain, further improving feature robustness. Our method not only preserves the semantic alignment benefits of CLIP but also enriches feature diversity and structural consistency across domains. Extensive experiments on the diverse weather-driving benchmark demonstrate that our approach outperforms the existing state-of-the-art method.

目标检测域泛化傅里叶变换vMF分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。