用量子思想提升分类体系扩展,让实体理解更精准。
QuanTaxo: A Quantum Approach to Self-Supervised Taxonomy Expansion
- 将实体编码到希尔伯特空间,模拟语义干扰效应。
- 在五个数据集上准确率提升12.3%,MRR提升11.2%。
- 适合需要动态更新知识体系的场景,如智能搜索与推荐。
分类体系是包含知识的层级图,可为多种网络应用提供重要洞察。然而,手动构建分类体系需大量人力,且随着网络内容爆炸式增长,现有分类体系容易过时,难以有效融入新信息。因此,亟需动态扩展分类体系以保持其相关性。现有方法多依赖经典词嵌入表示实体,但无法捕捉层级多义性——即实体的语义会随其在层级中的位置及上下文变化。为此,我们提出QuanTaxo,一种受量子启发的分类体系扩展框架,通过将实体编码至希尔伯特空间并建模其间的干涉效应,生成更丰富、上下文敏感的表示。在五个真实世界基准数据集上的全面实验表明,QuanTaxo显著优于九种经典嵌入基线模型,在准确率上提升12.3%,在平均倒数排名(MRR)上提升11.2%,在Wu & Palmer指标上提升6.9%。
原文摘要 · Abstract (English)
A taxonomy is a hierarchical graph containing knowledge to provide valuable insights for various web applications. However, the manual construction of taxonomies requires significant human effort. As web content continues to expand at an unprecedented pace, existing taxonomies risk becoming outdated, struggling to incorporate new and emerging information effectively. As a consequence, there is a growing need for dynamic taxonomy expansion to keep them relevant and up-to-date. Existing taxonomy expansion methods often rely on classical word embeddings to represent entities. However, these embeddings fall short of capturing hierarchical polysemy, where an entity's meaning can vary based on its position in the hierarchy and its surrounding context. To address this challenge, we introduce QuanTaxo, a quantum-inspired framework for taxonomy expansion that encodes entities in a Hilbert space and models interference effects between them, yielding richer, context-sensitive representations. Comprehensive experiments on five real-world benchmark datasets show that QuanTaxo significantly outperforms classical embedding models, achieving substantial improvements of 12.3% in accuracy, 11.2% in Mean Reciprocal Rank (MRR), and 6.9% in Wu & Palmer (Wu&P) metrics across nine classical embedding-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。