通过频域与像素域联合增强,提升模型在分布外数据上的鲁棒性。
Improving Out-of-Domain Robustness with Targeted Augmentation in Frequency and Pixel Spaces
- 融合源图幅度谱与目标图像素内容生成新样本
- 在四个跨领域基准上显著优于通用与专用增强方法
- 无需领域知识,适用于视觉、医疗、音频等多场景
在领域自适应设置下,带标签的源数据与无标签的目标数据来自不同分布,如何提升模型在分布外(OOD)场景下的鲁棒性是真实应用中的关键挑战。现有方法多依赖数据增强,但通用增强在面对目标域分布偏移时提升有限。虽有针对特定数据集的针对性增强可缓解此问题,但通常需专家经验与大量前期分析。为此,我们提出频率-像素连接(Frequency-Pixel Connect)框架,通过在频域与像素域同时引入针对性增强,提升领域自适应中的OOD鲁棒性。具体而言,将源图像的幅度谱与目标图像的像素内容混合,生成既引入领域多样性又保持源图像语义结构的增强样本。相比以往仅限于像素空间且依赖数据集特性的方法,本方法为数据集无关设计,具备更广适用性,超越自然图像范畴。我们进一步通过评估同类别跨域样本的连通性与异类样本的分离性,验证其有效性。实验表明,该方法在涵盖视觉、医学、音频和天文领域的四个真实世界基准上均显著优于现有通用与专用增强方法。
原文摘要 · Abstract (English)
Out-of-domain (OOD) robustness under domain adaptation settings, where labeled source data and unlabeled target data come from different distributions, is a key challenge in real-world applications. A common approach to improving OOD robustness is through data augmentations. However, in real-world scenarios, models trained with generic augmentations can only improve marginally when generalized under distribution shifts toward unlabeled target domains. While dataset-specific targeted augmentations can address this issue, they typically require expert knowledge and extensive prior data analysis to identify the nature of the datasets and domain shift. To address these challenges, we propose Frequency-Pixel Connect, a domain-adaptation framework that enhances OOD robustness by introducing a targeted augmentation in both the frequency space and pixel space. Specifically, we mix the amplitude spectrum and pixel content of a source image and a target image to generate augmented samples that introduce domain diversity while preserving the semantic structure of the source image. Unlike previous targeted augmentation methods that are both dataset-specific and limited to the pixel space, Frequency-Pixel Connect is dataset-agnostic, enabling broader and more flexible applicability beyond natural image datasets. We further analyze the effectiveness of Frequency-Pixel Connect by evaluating the performance of our method connecting same-class cross-domain samples while separating different-class examples. We demonstrate that Frequency-Pixel Connect significantly improves cross-domain connectivity and outperforms previous generic methods on four diverse real-world benchmarks across vision, medical, audio, and astronomical domains, and it also outperforms other dataset-specific targeted augmentation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。