用轻量大模型从对话中提取用户偏好三元组,构建隐私保护推荐知识图谱
From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems
- 用Qwen和Gemma模型从对话文本中提取带Wikidata标识的结构化三元组
- 部分模型在三元组提取和推荐任务中均表现良好,且下游性能与抽取能力正相关
- 为隐私友好型个性化推荐提供可复现的构建方案,适合关注数据安全的研究者
个人知识图谱(PKGs)为建模用户偏好提供了隐私保护框架,但如何从非结构化、分散的对话数据中构建仍具挑战。本文提出一种可复现的流水线,利用轻量级大语言模型将对话中的“字符串”转化为语义“实体”,从中提取符合RDF规范、链接至Wikidata标识的用户偏好三元组,用于构建个人知识图谱。评估了基于Qwen和Gemma的模型在对话数据中提取三元组的能力,并检验生成图谱在下游推荐任务中的实用性。结果表明,某些模型在三元组提取与推荐性能之间表现出良好的一致性,其下游表现与其抽取能力成比例。该方法为隐私敏感场景下的个性化推荐提供了可行路径。
原文摘要 · Abstract (English)
Personal Knowledge Graphs (PKGs) offer a privacy-preserving framework for modeling user preferences, yet constructing them from unstructured, decentralized conversational data remains a challenge. This paper bridges the gap between conversational "strings" and semantic "things" by presenting a reproducible pipeline for extracting structured user-preference triples using lightweight Large Language Models. We evaluate Qwen- and Gemma-based models on their ability to extract RDF-compliant triples linked to Wikidata identifiers from conversational data for PKG construction. Our evaluation assesses both the semantic extraction fidelity and the utility of the resulting graphs in a downstream recommendation task. We found that certain models performed well and had proportionally high downstream performance relative to their triple extraction performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。