arXiv:2508.17663cs.LGcs.IR2025-08中稿 · International Jour…

将异质数据的共现关系映射到二维空间,直观展现跨领域不对称关联。

Heterogeneous co-occurrence embedding for visual information exploration

  • 通过最大化互信息,将异质元素嵌入双维隐空间。
  • 支持三域以上扩展,用总相关性统一建模多域依赖。
  • 基于条件概率着色,交互式探索跨域关系,适合数据分析场景。

本文提出一种用于共现数据的嵌入方法,旨在支持视觉信息探索。考虑异质领域间元素对的共现概率测量场景。该方法将异质元素映射至对应的二维隐空间,实现领域间非对称关系的可视化。核心思想是通过最大化互信息来嵌入元素,尽可能保留原始依赖结构。该方法可自然推广至三域及以上情形,使用总相关性(total correlation)作为互信息的广义形式。针对跨域分析,还提出一种基于条件概率为隐空间着色的可视化方法,支持用户交互式探索非对称关系。通过在形容词-名词数据集、NeurIPS数据集以及主语-谓语-宾语数据集上的应用,展示了该方法在域内与跨域分析中的有效性。

原文摘要 · Abstract (English)

This paper proposes an embedding method for co-occurrence data aimed at visual information exploration. We consider cases where co-occurrence probabilities are measured between pairs of elements from heterogeneous domains. The proposed method maps these heterogeneous elements into corresponding two-dimensional latent spaces, enabling visualization of asymmetric relationships between the domains. The key idea is to embed the elements in a way that maximizes their mutual information, thereby preserving the original dependency structure as much as possible. This approach can be naturally extended to cases involving three or more domains, using a generalization of mutual information known as total correlation. For inter-domain analysis, we also propose a visualization method that assigns colors to the latent spaces based on conditional probabilities, allowing users to explore asymmetric relationships interactively. We demonstrate the utility of the method through applications to an adjective-noun dataset, the NeurIPS dataset, and a subject-verb-object dataset, showcasing both intra- and inter-domain analysis.

嵌入模型可视化异质数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。