arXiv:2606.10358cs.LGcs.AI2026-06

用知识图谱增强稀疏离散数据的贝叶斯网络结构学习,提升建模精度。

KG-SoftMAP: Soft Knowledge-Graph Priors for Bayesian Network Structure Learning from Sparse Discrete Data

  • 将加权有向知识图谱转为置信度先验,融合到BDeu评分中
  • 在极低观测率(ρ=0.05)下仍能恢复部分有向结构,F1达0.19-0.32
  • 适合教育等数据稀疏领域,兼具可解释性与预测性能

从稀疏离散数据中学习贝叶斯网络(BN)结构极具挑战:当每条实例仅记录少数变量时,多数变量对缺乏联合观测,导致可靠评分困难,纯数据方法难以恢复结构。通常可获得不完美的领域知识,以加权有向知识图谱(KG)形式表达。KG-SoftMAP将此类KG编码为有限强度、置信度加权的边先验,并在最大后验(MAP)目标中将其以logit形式加入BDeu评分。在信息丰富但不完美KG的支持下,即便在极低观测率ρ=0.05时,仍能恢复部分有向结构,跨基准测试的定向F1(DF1)为0.19–0.32;在更高观测率下,ρ=0.20时达到0.44–0.66,ρ=0.40时为0.46–0.64。若无KG先验,对应平均DF1仅为0.00、0.19、0.21。通过污染、移除或模糊KG信号的应力测试,以及对大模型提取图谱的验证,表明恢复效果随KG质量升降。在三个无真实DAG的现实教育数据集上评估预测、校准和KG一致性:在短答案反馈(SAF)数据集上,KG-SoftMAP+VE实现失败类F1 0.75(逻辑回归为0.78),同时提供可解释概念图、校准的失败概率及基于部分观测概念证据的后验查询。其余数据集显示:弱启发式KG信号不影响预测,而独立专家本体则使学习图更贴近专家关联。

原文摘要 · Abstract (English)

Learning Bayesian network (BN) structure from sparse discrete data is hard: when each instance records only a few variables, most variable pairs lack the joint observations needed for reliable scoring, and data-only methods recover little structure. Imperfect domain knowledge, expressible as a weighted directed knowledge graph (KG), is often available. KG-SoftMAP encodes such a KG as a finite-strength, confidence-weighted edge prior and maximizes a MAP objective that adds this logit-form prior to the BDeu score. With an informative but imperfect KG, KG-SoftMAP recovers partial directed structure even at observation rate rho=0.05, with directed F1 (DF1) of 0.19-0.32 across benchmarks. At higher observation rates within this sparse grid, DF1 reaches 0.44-0.66 at rho=0.20 and 0.46-0.64 at rho=0.40. Across the same three rates, KG-SoftMAP without the KG prior averages DF1 0.00, 0.19, and 0.21. Stress tests that corrupt, remove, or blur the KG signal, together with checks on LLM-extracted graphs beyond canonical benchmarks, show that recovery rises and falls with KG quality. On three real sparse educational datasets without ground-truth DAGs, we evaluate prediction, calibration, and KG-consistency. On Short Answer Feedback (SAF), KG-SoftMAP+VE reaches Fail-class F1 0.75 versus 0.78 for logistic regression while also providing an inspectable concept graph, calibrated Fail probabilities, and posterior queries from partially observed concept evidence. The remaining datasets sharpen the operating picture: weak heuristic KG signal leaves prediction unchanged, while an independent expert ontology moves the learned graph toward expert relatedness.

贝叶斯网络知识图谱稀疏数据可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。