用拓扑方法分析句子相似性,发现能有效捕捉语义关系。
Exploring Dowker Homology for Sentence Similarity

- 将句中词嵌入视为潜空间中的点云对,用Dowker同调分析相似性
- 通过回归验证其特征与真实相似度得分相关,相关性显著
- 可生成简洁数值摘要,适合模型可视化与对比分析
Dowker同调是一种可用于分析共同空间中两组点云相对位置的拓扑工具。本文探究其是否能捕捉句子相似性信息,将构成句子对的词嵌入视为变换器模型潜空间中的两组点云,分别使用经过和未经过句子相似性微调的模型进行实验。结果表明,基于回归分析,Dowker同调特征与真实相似度得分存在显著相关性,且可用于相似性数据与模型的可视化。为进一步提升实用性,我们从中推导出单数值摘要,期望直接反映句子相似性。这些摘要表现良好,但尚未超越基于标准池化方法的成熟句子相似性度量。
原文摘要 · Abstract (English)
Dowker homology is a topological tool that may be used to analyze the relative position of two point clouds living in a common space. We investigate whether Dowker homology captures sentence similarity information by treating the embeddings of the tokens that constitute a sentence pair as a pair of point clouds in the latent space of a transformer model, using both models that have and have not been fine-tuned for sentence similarity. We find that Dowker homology captures sentence similarity information, as measured by regressing Dowker homology features onto ground-truth similarity scores, and that it can be used for visual inspection of similarity data and models. In an attempt to make Dowker homology readily applicable, we derive from it single-number summaries that we expect to capture sentence similarity directly. These turn out to work reasonably well, but without outperforming standard sentence similarity measures based on established pooling methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。