arXiv:2608.24534cs.AI2026-08

用生理知识图谱提升临床大模型安全性,让生成建议更符合人体规律。

Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models

  • 训练时融合生理知识图谱与大模型,通过多跳路径和药物相互作用约束评分
  • 安全得分提升至90.8%,医生评估错误率降至5.1%,优于GPT-4
  • 适合医疗AI安全研究者,尤其关注可验证的临床推理系统

临床大模型生成的建议可能看似合理却违背生理规律。本文提出神经符号对齐框架,将70亿参数临床大模型与基于HGNN的生理世界模型结合,利用包含84.7万个节点的生物医学知识图谱进行训练。候选回答通过稳态约束、多跳路径合理性及药物相互作用惩罚打分,驱动迭代式策略优化。在2500场景的临床安全基准(CSB)上,相对ORPO方法,CSS从69.5%提升至90.8%(+21.3个百分点),医生盲评错误率由14.1%降至5.1%,DID从72.8%升至91.6%。独立规则引擎评分达86.4%(+21.2个百分点),与CSS相关性r=0.97。即使参数量仅为GPT-4的1/10,仍全面超越其表现;且优于推理时自纠错流程(+11.4个百分点CSS)。在模拟电子病历噪声下仍保持84.2%的CSS。消融实验显示,HGNN评分贡献-16.2个百分点,迭代训练贡献-11.5个百分点。基于200名医生标注的PhysioScore校准得ECE=0.038,kappa=0.91。结论:训练期生理知识锚定可带来可量化、可验证的安全提升,但需真实临床数据验证其部署效果。

原文摘要 · Abstract (English)

Clinical LLMs can generate recommendations that are factually plausible yet physiologically unsafe. We investigate whether safety alignment can be improved by grounding preference optimization in structured physiological knowledge rather than text-only supervision. Methods: We propose Neurosymbolic Alignment, a training-time framework that couples a 7B clinical LLM with an HGNN-based Physiological World Model over an 847K-node biomedical knowledge graph. Candidate responses are scored using homeostatic constraints, multi-hop path plausibility, and drug-interaction penalties, and the resulting rankings drive iterative on-policy ORPO updates. Evaluation is performed on the Clinical Safety Benchmark (CSB), a 2,500-scenario benchmark for physiological constraint violations in generative clinical reasoning. Results: Relative to ORPO, the proposed method improves CSS from 69.5% to 90.8% (+21.3 pp), reduces physician-evaluated HR from 14.1% to 5.1% on the blinded subset, and improves DID from 72.8% to 91.6%. These gains are corroborated by an HGNN-independent Rule-Engine Safety Score (RSS: 86.4%, +21.2 pp over ORPO; r=0.97 concordance with CSS). The method also exceeds GPT-4 (5-shot) on all safety metrics despite a 10x parameter disadvantage, and outperforms an inference-time self-correction pipeline (SFT+SelfCorrect) by 11.4 pp CSS. Under synthetic EHR-style noise, 84.2% CSS is retained. Ablation analysis shows that HGNN scoring (-16.2 pp) and iterative training (-11.5 pp) are the dominant contributors. PhysioScore calibration against 200 clinician labels yielded ECE = 0.038 and kappa = 0.91. Conclusion: Training-time physiological grounding produces measurable and independently verifiable safety improvements in open-weight clinical LLMs under controlled evaluation. External validation on real clinical data is required to determine whether these gains transfer to deployment settings

临床AI安全对齐知识图谱大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。