发现嵌入向量中语义成分的高阶关联,揭示词语间深层语义联系。
Understanding Higher-Order Correlations Among Semantic Components in Embeddings
- 用高阶相关性量化语义成分间的非独立性
- 高阶相关性强时,两成分共享大量同义词
- 通过最大生成树可视化语义成分关联结构
独立成分分析(ICA)可提取嵌入向量中可解释的语义成分。尽管ICA理论假设嵌入可线性分解为独立成分,但真实数据通常不满足此条件,导致估计成分间仍存在非独立性,而ICA无法消除。本文通过高阶相关性量化这些非独立性,发现当两个成分的高阶相关性较大时,表明二者具有强烈语义关联,且许多词汇同时与两者意义相近。整体非独立结构通过语义成分的最大生成树进行可视化,为基于ICA的嵌入分析提供了更深入的理解。
原文摘要 · Abstract (English)
Independent Component Analysis (ICA) offers interpretable semantic components of embeddings. While ICA theory assumes that embeddings can be linearly decomposed into independent components, real-world data often do not satisfy this assumption. Consequently, non-independencies remain between the estimated components, which ICA cannot eliminate. We quantified these non-independencies using higher-order correlations and demonstrated that when the higher-order correlation between two components is large, it indicates a strong semantic association between them, along with many words sharing common meanings with both components. The entire structure of non-independencies was visualized using a maximum spanning tree of semantic components. These findings provide deeper insights into embeddings through ICA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。