arXiv:2604.03953cs.CVcs.LG2026-04

提出CM-GLasso模型,分离多模态数据中的共性与特异性结构。

Multimodal Structure Learning: Disentangling Shared and Specific Topology via Cross-Modal Graphical Lasso

  • 通过跨模态注意力蒸馏,将高维特征压缩为语义节点。
  • 联合优化共享与特定结构,实现精度矩阵的同步解耦。
  • 在8个基准上超越现有方法,适用于医学与自然图像任务。

学习可解释的多模态表示依赖于发现异构特征间的条件依赖关系。然而,如图稀疏估计技术(如Graphical Lasso, GLasso)在视觉-语言领域面临高维噪声、模态错位以及共性与类别特异性拓扑混淆等瓶颈。本文提出跨模态图稀疏学习(CM-GLasso),通过新颖的文本-视觉对齐策略与统一的视觉-语言编码器,严格对齐多模态特征至共享隐空间。引入跨注意力蒸馏机制,将高维图像块浓缩为显式语义节点,自然提取空间感知的跨模态先验。进一步将定制化GLasso估计与共性-特异性结构学习(CSSL)统一为联合目标,采用交替方向乘子法(ADMM)优化,保证不变与类别特异性精度矩阵的同步解耦,避免多步误差累积。在涵盖自然与医学领域的8个基准上进行大量实验,结果表明CM-GLasso在生成分类与密集语义分割任务中达到新SOTA水平。

原文摘要 · Abstract (English)

Learning interpretable multimodal representations inherently relies on uncovering the conditional dependencies between heterogeneous features. However, sparse graph estimation techniques, such as Graphical Lasso (GLasso), to visual-linguistic domains is severely bottlenecked by high-dimensional noise, modality misalignment, and the confounding of shared versus category-specific topologies. In this paper, we propose Cross-Modal Graphical Lasso (CM-GLasso) that overcomes these fundamental limitations. By coupling a novel text-visualization strategy with a unified vision-language encoder, we strictly align multimodal features into a shared latent space. We introduce a cross-attention distillation mechanism that condenses high-dimensional patches into explicit semantic nodes, naturally extracting spatial-aware cross-modal priors. Furthermore, we unify tailored GLasso estimation and Common-Specific Structure Learning (CSSL) into a joint objective optimized via the Alternating Direction Method of Multiplier (ADMM). This formulation guarantees the simultaneous disentanglement of invariant and class-specific precision matrices without multi-step error accumulation. Extensive experiments across eight benchmarks covering both natural and medical domains demonstrate that CM-GLasso establishes a new state-of-the-art in generative classification and dense semantic segmentation tasks.

多模态图学习结构解耦视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。