arXiv:2603.21139cs.IRcs.LG2026-03

用知识图谱和用户画像实现更精准的XML文档搜索

Ontology-driven personalized information retrieval for XML documents

  • 将文档、查询和用户画像转为加权概念向量,融合领域本体
  • 相比关键词匹配,精确率与召回率均显著提升
  • 适合需要个性化检索的领域专家或复杂信息查询者

本文针对半结构化XML文档的信息检索难题,提出一种基于本体与用户画像的个性化检索框架。传统检索系统忽视用户差异,对同一查询返回相同结果。本文引入领域本体与用户档案,将文档、查询及用户画像表示为加权概念向量。本体通过概念加权机制突出层次结构中较低层级的特定概念,以提供更精确的信息。利用语义相似度度量捕捉概念间关系,超越关键词匹配,实现用户、查询与文档间的细粒度个性化匹配。实验表明,结合本体与用户画像的方案在精度和召回率上均优于传统关键词方法。该框架显著提升了XML搜索结果的相关性与适应性,支持更以用户为中心的检索体验。

原文摘要 · Abstract (English)

This paper addresses the challenge of improving information retrieval from semi-structured eXtensible Markup Language (XML) documents. Traditional information retrieval systems (IRS) often overlook user-specific needs and return identical results for the same query, despite differences in users' knowledge, preferences, and objectives. We integrate external semantic resources, namely a domain ontology and user profiles, into the retrieval process. Documents, queries, and user profiles are represented as vectors of weighted concepts. The ontology applies a concept-weighting mechanism that emphasizes highly specific concepts, as lower-level nodes in the hierarchy provide more precise and targeted information. Relevance is assessed using semantic similarity measures that capture conceptual relationships beyond keyword matching, enabling personalized and fine-grained matching among user profiles, queries, and documents. Experimental results show that combining ontologies with user profiles improves retrieval effectiveness, achieving higher precision and recall than keyword-based approaches. Overall, the proposed framework enhances the relevance and adaptability of XML search results, supporting more user-centered retrieval.

信息检索本体个性化XML

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。