构建保持子空间结构的稀疏注意力图,实现异构多视图数据的语义对齐。
Learning Subspace-Preserving Sparse Attention Graphs from Heterogeneous Multiview Data

- 用双线性注意力捕捉高维特征间的非对称相似性,突破传统对称瓶颈。
- 动态稀疏门控自适应调节邻居贡献,生成每视图特有的稀疏图。
- 理论证明可微稀疏注意力与概率单纯形约束的关联,适合多视图表征学习。
通过不同架构的预训练模型从大规模无标签数据中提取的高维特征称为异构多视图数据。现有大多数无监督迁移学习方法在利用多视图互补信息时,难以准确恢复内在子空间结构。因此,关键挑战在于构建能保持底层子空间结构的稀疏相似性图,以实现异构视图间的语义对齐。本文提出稀疏注意力图学习(SAGL)方法,从异构多视图数据中学习保持子空间结构的稀疏注意力图。具体地,引入双线性注意力分解机制,捕捉高维特征间的非对称相似性,打破传统表示学习中的对称性瓶颈。动态稀疏门控机制预测特征特定的压缩因子,自适应控制邻居的拓扑贡献。此外,采用α-entmax结构化稀疏投影,为各视图生成保持子空间结构的稀疏注意力图。SAGL利用这些视图特异性图进行稀疏信息聚合,获得用于多视图学习任务的判别性表示。进一步提供严格的理论分析,建立可微稀疏注意力与概率单纯形约束之间的桥梁。在多个基准数据集上的大量实验表明,SAGL持续优于当前最先进的无监督迁移学习方法。
原文摘要 · Abstract (English)
The high-dimensional features extracted from large-scale unlabeled data via various pretrained models with diverse architectures are referred to as heterogeneous multiview data. Most existing unsupervised transfer learning methods fail to faithfully recover intrinsic subspace structures when exploiting complementary information across multiple views. Therefore, a fundamental challenge involves constructing sparse similarity graphs that preserve these underlying subspace structures for achieving semantic alignment across heterogeneous views. In this paper, we propose a sparse attention graph learning (SAGL) method that learns subspace-preserving sparse attention graphs from heterogeneous multiview data. Specifically, we introduce a bilinear attention factorization scheme to capture asymmetric similarities among the high-dimensional features, which breaks the symmetry bottleneck that is inherent in the traditional representation learning techniques. A dynamic sparsity gating mechanism then predicts a feature-specific compression factor for adaptively controlling the topological contributions of neighbors. Furthermore, we employ a structured sparse projection via $α$-entmax to generate subspace-preserving sparse attention graphs for individual views. SAGL leverages these view-specific graphs to conduct sparse information aggregation, yielding discriminative representations for multiview learning tasks. In addition, we provide a rigorous theoretical analysis that bridges differentiable sparse attention and probability simplex constraints. Extensive experiments conducted on multiple benchmark datasets demonstrate that SAGL consistently outperforms the state-of-the-art unsupervised transfer learning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。