用文本引导视觉依赖图学习,让模型自动发现跨模态的隐藏关联。
Text-Guided Visual Dependency Graph Learning with Cross-Modal Attention Priors

- 通过文本转图像+跨模态注意力,生成可解释的视觉依赖结构先验
- 在VOC和ADE20K上达91.97%分类准确率、74.75%和64.01%分割mIoU
- 适合需要可解释性与轻量级推理的多模态理解任务
从多模态视觉-语言特征中估计可解释的条件依赖结构仍处于未探索状态。我们提出CM-GLasso(跨模态图lasso)框架,连接视觉语言表征学习与稀疏高斯图模型。该方法包含三个关键组件:(i) 文本可视化策略,将类别属性描述转为图像,并通过与自然图像相同的SigLIP-2视觉编码器处理,获得共享特征空间中的补丁级注意力足迹;(ii) 跨注意力蒸馏机制,将高维补丁压缩为少量语义图节点,其注意力相似性提供非均匀L1正则化的跨模态结构先验;(iii) 联合ADMM公式,统一优化共享与类别特异性精度成分,避免先独立建模再分解的步骤。学习到的稀疏图拓扑直接支持无参数的基于精度的分类规则和轻量级拓扑感知分割头。在八个基准上的大量实验表明,CM-GLasso在匹配控制协议下优于或媲美强基线,在VOC(74.75% mIoU)和ADE20K(64.01% mIoU)上达到最高平均分类准确率(91.97%),同时生成具有共性-特性分解的显式稀疏条件依赖图。
原文摘要 · Abstract (English)
Estimating interpretable conditional-dependence structures from multimodal visual-linguistic features remains largely unexplored. We propose CM-GLasso (Cross-Modal Graphical Lasso), a framework that bridges vision-language representation learning and sparse Gaussian Graphical Models. CM-GLasso introduces three key components: (i) a text visualization strategy that renders class-attribute descriptions as images and processes them through the same SigLIP-2 vision encoder as natural images, yielding prototype-indexed patch-level attention footprints in a shared feature coordinate system; (ii) a cross-attention distillation mechanism that condenses high-dimensional patches into a small set of semantic graph nodes, whose attention-footprint similarities yield cross-modal structural priors for non-uniform L1 penalization; (iii) a joint ADMM formulation that estimates shared and class-specific precision components within a single convex objective, avoiding the need to first estimate and then decompose separate class-wise graphs. The learned sparse graph topologies directly support a parameter-free, precision-based classification rule and a lightweight topology-aware segmentation head. Extensive experiments on eight benchmarks demonstrate that CM-GLasso achieves competitive or superior performance compared with strong feature-based and task-specific baselines. Under the matched controlled protocol, it attains the highest average classification accuracy (91.97%) and the highest segmentation mIoU among the controlled baselines on VOC (74.75%) and ADE20K (64.01%), while also yielding explicit sparse conditional-dependence graphs with common-specific decomposition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。