arXiv:2602.04328cs.CV2026-02被引 1

跨异构视图学习不变特征,无需标签也能提升模型泛化能力

Multiview Self-Representation Learning across Heterogeneous Views

  • 用自表示机制融合不同预训练模型的特征输出
  • 在多个视觉数据集上优于现有方法,显著提升特征一致性
  • 适合无监督迁移学习、多模型特征对齐场景

同一样本经不同预训练模型生成的特征因预训练目标或结构差异,常呈现明显不同的分布。如何在无标签大规模视觉数据上,以全无监督方式从多种预训练模型中学习不变表示,仍是重大挑战。本文提出多视图自表示学习(MSRL)方法,通过利用异构视图下特征的自表示特性来学习不变表示。特征由多种预训练模型通过迁移学习从大规模无标签视觉数据中提取,称为异构多视图数据。每个视图在对应冻结的预训练主干网络上叠加一个独立线性模型,并引入基于自表示学习的信息传递机制,支持线性模型输出的特征聚合。同时,设计了分配概率分布一致性方案,通过挖掘不同视图间的互补信息,引导多视图自表示学习,从而强制不同线性模型间的表示不变性。此外,本文提供了信息传递机制、分配概率分布一致性和增量视图的理论分析。大量实验在多个基准视觉数据集上验证,所提方法持续优于多个前沿方法。

原文摘要 · Abstract (English)

Features of the same sample generated by different pretrained models often exhibit inherently distinct feature distributions because of discrepancies in the model pretraining objectives or architectures. Learning invariant representations from large-scale unlabeled visual data with various pretrained models in a fully unsupervised transfer manner remains a significant challenge. In this paper, we propose a multiview self-representation learning (MSRL) method in which invariant representations are learned by exploiting the self-representation property of features across heterogeneous views. The features are derived from large-scale unlabeled visual data through transfer learning with various pretrained models and are referred to as heterogeneous multiview data. An individual linear model is stacked on top of its corresponding frozen pretrained backbone. We introduce an information-passing mechanism that relies on self-representation learning to support feature aggregation over the outputs of the linear model. Moreover, an assignment probability distribution consistency scheme is presented to guide multiview self-representation learning by exploiting complementary information across different views. Consequently, representation invariance across different linear models is enforced through this scheme. In addition, we provide a theoretical analysis of the information-passing mechanism, the assignment probability distribution consistency and the incremental views. Extensive experiments with multiple benchmark visual datasets demonstrate that the proposed MSRL method consistently outperforms several state-of-the-art approaches.

无监督学习多视图学习特征对齐自表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。