arXiv:2607.27660cs.LGcs.AI2026-07

揭示子模信息度量如何影响表示学习中的类内方差与类间分离,为选择目标函数提供理论依据。

Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective

  • 从方差与分离角度统一分析子模信息度量的几何统计特性。
  • 不同度量分别对应类内方差、协方差体积和稀有类分离等关键结构。
  • 适用于需要理解目标函数设计原理的研究者,尤其在不平衡多模态学习中。

子模信息度量(SIMs)近年来成为表示学习和多模态学习的强大框架。特别是SCORE框架表明,SIMs可作为监督对比学习的有效目标函数。尽管其经验表现优异,但不同子模信息度量所诱导的几何与统计特性仍不明确。本文建立了一个统一的理论框架,将SIMs与经典表示学习及统计模式识别概念相联系。我们发现:总信息(TI)目标刻画类内结构——图割TI恢复类内方差,LogDet TI恢复广义方差与协方差体积,设施选址TI实现对稀有和易混淆类的不平衡感知分离。互信息(MI)目标则捕捉互补的类间结构——图割MI关联中心点分离与费舍尔判别,LogDet MI通过马氏距离体现协方差感知分离,设施选址MI度量最近模式的表征重叠。通过受控合成实验,独立调节方差、协方差、类别不平衡、类间分离与多模态重叠,实证结果与理论预测高度一致。本工作首次提供了对子模信息度量的统一几何与统计理解,并为设计与选择基于SIM的目标函数提供原则性指导。

原文摘要 · Abstract (English)

Submodular Information Measures (SIMs) have recently emerged as a powerful framework for representation learning and multimodal learning. In particular, the SCORE framework~\cite{majee2024score} demonstrated that SIMs can serve as effective objectives for supervised contrastive learning. Despite their empirical success, however, the geometric and statistical properties induced by different submodular information measures remain poorly understood. In this work, we develop a unified theoretical framework connecting SIMs to classical concepts in representation learning and statistical pattern recognition. We show that Total Information (TI) objectives characterize intra-class structure: Graph Cut TI recovers within-class variance, LogDet TI recovers generalized variance and covariance volume, and Facility Location TI induces imbalance-aware separation that emphasizes rare and confusable classes. We further show that Mutual Information (MI) objectives capture complementary notions of inter-class structure: Graph Cut MI is closely related to centroid separation and Fisher-style discrimination, LogDet MI captures covariance-aware separation through Mahalanobis distance, and Facility Location MI measures nearest-mode representational overlap. We validate these theoretical characterizations using controlled synthetic experiments that independently vary variance, covariance, class imbalance, class separation, and multimodal overlap. Across all settings, the empirical behavior closely matches the proposed theory. Our results provide the first unified geometric and statistical understanding of submodular information measures and offer principled guidance for selecting and designing SIM-based objectives for representation learning.

表示学习子模优化信息度量理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。