通过分布式的联合潜在空间建模,提升多视角表征学习的泛化能力。
Multiview Representation Learning via Distributed Joint Latent Space Structuring
- 基于最小描述长度构建跨视角联合潜在变量的泛化界。
- 发现冗余表征有助于提升性能,且视图间相关性能收紧边界。
- 提出可分布式学习的高斯混合先验,捕捉视图间依赖关系。
我们研究分布式多视角表示学习,即K个客户端各自观测一个不同但可能统计相关的视图。客户端从各自视图中独立提取局部表示,再由中心解码器用于联合目标估计。核心难点在于客户端无法通信,必须自主决定编码内容。本文从泛化误差角度分析这一协调问题。针对分类与回归任务,我们推导出以所有视图在训练和测试数据上的联合潜在变量的最小描述长度(MDL)表示的新型泛化界。该结构感知界表明,提取表示间的统计相关性可收紧边界,为跨视图特征对齐的实证优势提供理论支持。反直觉的是,研究结果暗示编码器可能受益于提取冗余表示。受此启发,我们引入一种数据依赖的高斯乘积混合先验,可完全分布式学习与应用。该多视角先验的联合结构捕捉了通常被仅边缘化方法忽略的视图间依赖。在多个数据集、编码器架构、视图数量及失真设置下的全面实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
We study distributed multiview representation learning, a problem in which $K$ clients each observe a distinct but possibly statistically correlated view. The clients independently extract local representations from their views, which are then used by a central decoder for joint target estimation. One central difficulty is that, since the clients are not allowed to communicate with each other, they must autonomously decide what to encode. We study this coordination problem from a generalization error perspective. For both classification and regression tasks, we derive novel generalization bounds expressed in terms of the Minimum Description Length (MDL) of the joint latent variables across all views and across both training and test datasets. Our structure-aware bound reveals that statistical correlations among the extracted representations tighten the bound, providing theoretical grounding for the empirically observed benefits of cross-view feature alignment. Perhaps counterintuitively, our findings imply that encoders may benefit from extracting redundant representations. Motivated by these bounds, we introduce a data-dependent Gaussian product mixture prior that can be learned and applied in a fully distributed manner. The joint structure of this multiview prior captures inter-view dependencies that are typically discarded by marginal-only approaches. Comprehensive experiments across multiple datasets, encoder architectures, numbers of views, and distortion settings demonstrate the effectiveness of our proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。