arXiv:2607.00850cs.CVcs.LG2026-07中稿 · KDD

通过镜像融合注意力提升自监督学习对称性感知能力

Mirror-Fusion Attention for Reflection-Aware Self-Supervised Representation Learning

论文配图:Mirror-Fusion Attention for Reflection-Aware Self-Supervised Representation Learning
图 1 · 摘自论文原文
  • 引入镜像配对视图与轻量级融合注意力,动态交互对称区域
  • 在多个数据集上超越MoCo-v3等基线,反射鲁棒性显著提升
  • 仅增2.7%参数,适合医疗影像与人脸等对称数据任务

多数自监督学习方法追求变换下的不变性,但严格的翻转不变性会抑制医学影像和人脸等近似双侧对称数据中的左右对应信息。我们提出镜像融合增强的自监督学习(MFASSL),一种在不重设计主干网络的前提下注入软反射先验的视觉变换器框架。MFASSL构建沿估计对称轴对齐的镜像视图,并引入轻量级镜像融合注意力模块,在保留非对称线索的同时实现镜像区域间的自适应令牌级交互。基础自监督目标进一步耦合反射一致性与中间层令牌对齐损失。在CheXpert、BraTS、CelebA-HQ和WFLW数据集上,MFASSL在相同ViT-B/16设置下,优于MoCo-v3、DINO和MAE基线,提升了下游性能、校准度和反射鲁棒性。其增益强且稳定,仅增加约2.7%参数,优于近期等变自监督方法。结果表明,轻量级几何感知先验可有效补充基于不变性的自监督学习。

原文摘要 · Abstract (English)

Most self-supervised learning (SSL) methods encourage invariance across augmentations, but strict flip invariance can suppress informative left--right correspondences in approximately bilateral data such as medical images and human faces. We propose Mirror-Fusion-Augmented Self-Supervised Learning (MFASSL), a Vision Transformer framework that injects a soft reflection prior into standard SSL without redesigning the backbone. MFASSL constructs mirror-paired views aligned to an estimated symmetry axis and introduces a lightweight Mirror-Fusion Attention (MFA) module for adaptive token-level interaction between mirrored regions while preserving asymmetric cues. The base SSL objective is further coupled with reflection-consistency and mid-layer token-alignment losses. Across CheXpert, BraTS, CelebA-HQ, and WFLW, MFASSL improves downstream performance, calibration, and reflection robustness over MoCo-v3, DINO, and MAE baselines under matched ViT-B/16 settings. It also achieves stronger and more consistent gains than recent equivariant SSL approaches with only approximately 2.7\% additional parameters. These results show that lightweight geometry-aware priors can effectively complement invariance-based SSL.

自监督学习视觉变换器对称性感知镜像融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。