用领域知识增强语义检索,提升文档查找准确率
Enhancing Semantic Document Retrieval- Employing Group Steiner Tree Algorithm with Domain Knowledge Enrichment
- 基于分组斯坦纳树构建融合领域知识的语义检索算法
- 在真实查询上达到90%精度和82%准确率
- 适合需要高精度文档检索的应用场景
从具有不同特征的数据源中检索相关文档对文档检索系统构成重大挑战,尤其当需考虑数据与领域知识间的语义关系时。现有基于语义的检索系统(通常以开放资源生成的知识图谱和通用领域知识表示)虽具潜力,但因缺乏领域特定信息且依赖过时知识源,精度可能受限。本研究提出两项核心贡献:一是开发一种名为‘基于语义的概念检索-分组斯坦纳树’(SemDR)的通用算法,通过融入领域知识提升语义感知的知识表示与数据访问能力;二是将该算法应用于真实世界数据的文档检索系统中。为评估系统效能,研究使用包含170个真实搜索查询的基准进行测试,并经领域专家严格评估以确保结果有效性和准确性。实验结果表明,相比基线系统,该方法在精度和准确率上分别达到90%和82%,表现出显著提升。
原文摘要 · Abstract (English)
Retrieving pertinent documents from various data sources with diverse characteristics poses a significant challenge for Document Retrieval Systems. The complexity of this challenge is further compounded when accounting for the semantic relationship between data and domain knowledge. While existing retrieval systems using semantics (usually represented as Knowledge Graphs created from open-access resources and generic domain knowledge) hold promise in delivering relevant outcomes, their precision may be compromised due to the absence of domain-specific information and reliance on outdated knowledge sources. In this research, the primary focus is on two key contributions- a) the development of a versatile algorithm- 'Semantic-based Concept Retrieval using Group Steiner Tree' that incorporates domain information to enhance semantic-aware knowledge representation and data access, and b) the practical implementation of the proposed algorithm within a document retrieval system using real-world data. To assess the effectiveness of the SemDR system, research work conducts performance evaluations using a benchmark consisting of 170 real-world search queries. Rigorous evaluation and verification by domain experts are conducted to ensure the validity and accuracy of the results. The experimental findings demonstrate substantial advancements when compared to the baseline systems, with precision and accuracy achieving levels of 90% and 82% respectively, signifying promising improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。