用逻辑语法检测医疗数据错误,自动修复异常记录。
Logical Grammar Induction via Graph Kolmogorov Complexity: A Neuro-Symbolic Framework for Self-Healing Clinical Data Integrity
- 将病历视为受隐含逻辑规则约束的私有语言,用图复杂度建模。
- 在200万条数据上实现0.94的F1分数,误报率降低12%。
- 适合需要实时纠错的医疗信息系统,兼顾准确与可解释性。
医疗信息系统的可靠性常因人为录入错误而受损,现有统计异常检测方法难以区分真实临床极端值与数据错误。本文提出Logic-GNN,一种神经符号框架,将临床记录视为由潜在逻辑规则支配的结构化“私有语言”。通过结合时序图神经网络(TGNN)与图柯尔莫哥洛夫复杂度,推导出表示医疗交互底层逻辑的符号语法。我们将异常定义为导致临床图最小描述长度(MDL)显著增加的“语法违规”。在包含200万+记录的Sina System数据集上,Logic-GNN取得0.94的F1分数,相比最先进基线提升12%,能有效区分危及生命的医学异常与数据损坏。该方法引入自愈机制,可在实时医疗信息系统中建议逻辑修正以维护数据完整性。
原文摘要 · Abstract (English)
The reliability of Healthcare Information Systems (HIS) is frequently compromised by human-induced data entry errors, which existing statistical anomaly detection methods fail to distinguish from legitimate clinical extremes. This paper proposes Logic-GNN, a novel neuro-symbolic framework that treats clinical records as a structured ``private language'' governed by latent logical games. By integrating Temporal Graph Neural Networks (TGNN) with Graph Kolmogorov Complexity, we induce a symbolic grammar that represents the underlying logic of medical interactions. We define anomalies as ``grammatical violations'' that cause a significant expansion in the Minimum Description Length (MDL) of the clinical graph. Evaluated on the Sina System dataset (2M+ records), Logic-GNN achieves an F1-score of 0.94, outperforming state-of-the-art baselines by 12\% in distinguishing between life-threatening medical outliers and data corruption. Our approach introduces a self-healing mechanism that suggests logical corrections to maintain data integrity in real-time HIS environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。