arXiv:2512.21011cs.CV2025-12被引 1

通过结构感知掩码增强模型鲁棒性,避免关键信息丢失。

Granular Ball Guided Masking: Structure-aware Data Augmentation

  • 基于粒度球计算,分层自适应保留重要结构区域
  • 在多个基准上提升分类与图像篡改检测性能
  • 适用于CNN和ViT,无需修改模型架构

深度学习在计算机视觉中取得显著成功,但仍严重依赖大规模标注数据,在数据有限或分布变化时易过拟合。基于掩码的信息丢弃类数据增强可通过强制模型探索互补线索来提升鲁棒性,但现有方法常缺乏结构感知,可能丢弃关键语义。本文提出粒度球引导掩码(GBGM),一种基于粒度球计算(GBC)的结构感知增强策略。GBGM通过粗到细的层级掩码过程,自适应保留语义丰富、结构重要的区域,同时抑制冗余区域,生成既具代表性又具区分性的增强样本。在多个基准上的大量实验表明,GBGM不仅在图像分类和掩码图像重建任务中表现持续提升,还在图像篡改检测中验证了有效性,证明其在识别与取证场景下的通用性。该方法简单且模型无关,可无缝集成于CNN与视觉变换器,提供了一种实用的结构感知数据增强范式。

原文摘要 · Abstract (English)

Deep learning models have achieved remarkable success in computer vision but still rely heavily on large-scale labeled data and tend to overfit when data is limited or distributions shift. Data augmentation -- particularly mask-based information dropping -- can enhance robustness by forcing models to explore complementary cues; however, existing approaches often lack structural awareness and risk discarding essential semantics. We propose Granular Ball Guided Masking (GBGM), a structure-aware augmentation strategy guided by Granular Ball Computing (GBC). GBGM adaptively preserves semantically rich, structurally important regions while suppressing redundant areas through a coarse-to-fine hierarchical masking process, producing augmentations that are both representative and discriminative. Extensive experiments on multiple benchmarks demonstrate consistent improvements not only in image classification and masked image reconstruction, but also in image tampering detection, validating the effectiveness and generalization of GBGM across both recognition and forensic scenarios. Simple and model-agnostic, GBGM integrates seamlessly into CNNs and Vision Transformers, offering a practical paradigm for structure-aware data augmentation.

数据增强结构感知图像分类掩码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。