arXiv:2511.22150cs.LGcs.CL2025-11被引 3

用统一拓扑签名解析文本嵌入空间,预测检索效果。

From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures

  • 提出统一拓扑签名框架,综合多维度刻画嵌入空间结构。
  • 发现模型架构相似性会引发嵌入空间结构趋同。
  • 可准确预测文档可检索性,适合关注模型可解释性的研究者。

研究嵌入空间的组织方式不仅提升模型可解释性,还揭示影响下游任务性能的关键因素。本文对多种文本嵌入模型和数据集进行了广泛的拓扑与几何度量分析,发现这些度量具有高度冗余性,单一指标常无法有效区分嵌入空间。基于此,我们提出统一拓扑签名(UTS)框架,用于整体表征嵌入空间。实验表明,UTS可预测模型特有属性,并揭示由模型架构驱动的空间相似性。进一步验证显示,拓扑结构与排序有效性密切相关,且能准确预测文档可检索性。结果表明,理解与利用文本嵌入的几何特性,必须采用多属性融合的综合性视角。

原文摘要 · Abstract (English)

Studying how embeddings are organized in space not only enhances model interpretability but also uncovers factors that drive downstream task performance. In this paper, we present a comprehensive analysis of topological and geometric measures across a wide set of text embedding models and datasets. We find a high degree of redundancy among these measures and observe that individual metrics often fail to sufficiently differentiate embedding spaces. Building on these insights, we introduce Unified Topological Signatures (UTS), a holistic framework for characterizing embedding spaces. We show that UTS can predict model-specific properties and reveal similarities driven by model architecture. Further, we demonstrate the utility of our method by linking topological structure to ranking effectiveness and accurately predicting document retrievability. We find that a holistic, multi-attribute perspective is essential to understanding and leveraging the geometry of text embeddings.

嵌入空间拓扑分析可检索性模型解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。