用检索增强小模型+形式概念分析,让知识扩展更可信。
Verifiable Knowledge Expansion through Retrieval-Grounded Formal Concept Analysis

- 用形式概念分析生成知识推论,再由检索增强小模型验证或反驳。
- 在罕见共济失调数据上,关系F1达0.29-0.52,推论F1为0.22-0.30。
- 可追踪每条推论的对错与修正,适合医学等高可靠性知识构建场景。
本研究针对知识图谱构建中结构合理性问题,提出一种基于检索增强的小语言模型框架,结合形式概念分析(FCA)作为符号化验证环。从初始属性出发,FCA逐步生成蕴含关系,再由检索增强的轻量级语言模型(SLM)对每条推论进行验证或返回反例。该模型还支持属性关联判断、一致性检查和新属性提议,使被接受的推论、反例、矛盾和修正均具可追溯性。在基于Orphadata资源构建的罕见共济失调数据集上,10个种子属性的实验获得关系F1为0.29–0.52,闭包推论F1为0.22–0.30。更大的种子集合可评估更多推论并常提升推论性能。较低的推论分数反映对推导关系的严格检验标准,一条遗漏或错误关系可能影响多个推论判断。消融实验表明,在固定对象-属性设定下进行关联判断可提升闭包推论得分,但即使在已知候选集情况下,识别正向对象-属性对仍具挑战。
原文摘要 · Abstract (English)
Ontology construction requires deciding which objects, attributes, and structural relations should be accepted as valid knowledge. Language models can propose such structures from text, but their outputs can still be unsupported or inconsistent. This paper proposes a retrieval-augmented small language model (SLM) framework that uses formal concept analysis (FCA) as a symbolic verification loop for knowledge expansion. Starting from seed attributes, FCA proposes implications over a growing formal context. A retrieval-grounded SLM oracle then validates each implication or returns a counterexample. The oracle also supports incidence judgments, consistency checks, and attribute proposals, making accepted implications, counterexamples, contradictions, and corrections inspectable. In a rare ataxia setting constructed from Orphadata resources, retrieval-grounded 10-seed runs obtain relation F1 of 0.29-0.52 and closure-based implication F1 of 0.22-0.30. Larger seed sets increase the number of evaluated implications and often improve implication F1. The lower implication scores reflect a stricter evaluation of derived implications, where one missed or extra relation can affect several implication judgments. Ablations show that incidence judgments in a fixed object-attribute setting can improve closure-based implication scores. However, identifying positive object-attribute pairs remains difficult even when the candidate objects and attributes are fixed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。