arXiv:2603.08321cs.AI2026-03

用可追溯推理和知识图谱确保针灸诊疗AI安全可靠

CORE-Acu: Structured Reasoning Traces and Knowledge Graph Safety Verification for Acupuncture Clinical Decision Support

  • 构建结构化推理链,让中医思维过程可追踪可解释
  • 针灸安全知识图谱零违规,显著优于GPT-4o的8.5%违规率
  • 专为高风险术语设计损失函数,提升关键实体准确性

大型语言模型在临床决策支持中潜力巨大,但其黑箱特性——推理不可追溯、存在概率性幻觉——在要求高度可解释性和安全性的针灸领域构成严重挑战。为此,我们提出CORE-Acu,一种融合结构化思维链(S-CoT)与知识图谱(KG)安全验证的神经符号框架。首先,构建首个针灸结构化推理轨迹数据集及模式约束微调框架,通过显式因果链从证型识别到治则、方案再到穴位选择,将传统中医推理转化为可解释的生成约束,缓解基于LLM的临床决策支持的不透明性。其次,构建中医安全知识图谱,建立基于符号否决机制的‘生成-验证-修正’闭环推理系统,以确定性规则拦截幻觉并强制执行硬性安全边界。最后,提出词典匹配实体重加权损失(LMERL),通过自适应放大高频高危实体在微调中的梯度贡献,纠正通用优化中因频率-重要性失配导致的术语漂移。在1000例保留病例上的实验表明,CORE-Acu在实体保真度和推理质量上均表现优异,最关键的是,在相同规则下,其安全违规率为0/1000(95%置信区间:0–0.37%),而GPT-4o为8.5%。这些结果确立了CORE-Acu作为针灸临床决策支持中兼具可审计推理与严格安全合规的稳健神经符号框架。

原文摘要 · Abstract (English)

Large language models (LLMs) show significant potential for clinical decision support (CDS), yet their black-box nature -- characterized by untraceable reasoning and probabilistic hallucinations -- poses severe challenges in acupuncture, a field demanding rigorous interpretability and safety. To address this, we propose CORE-Acu, a neuro-symbolic framework for acupuncture clinical decision support that integrates Structured Chain-of-Thought (S-CoT) with knowledge graph (KG) safety verification. First, we construct the first acupuncture Structured Reasoning Trace dataset and a schema-constrained fine-tuning framework. By enforcing an explicit causal chain from pattern identification to treatment principles, treatment plans, and acupoint selection, we transform implicit Traditional Chinese Medicine (TCM) reasoning into interpretable generation constraints, mitigating the opacity of LLM-based CDS. Furthermore, we construct a TCM safety knowledge graph and establish a ``Generate--Verify--Revise'' closed-loop inference system based on a Symbolic Veto Mechanism, employing deterministic rules to intercept hallucinations and enforce hard safety boundaries. Finally, we introduce the Lexicon-Matched Entity-Reweighted Loss (LMERL), which corrects terminology drift caused by the frequency--importance mismatch in general optimization by adaptively amplifying gradient contributions of high-risk entities during fine-tuning. Experiments on 1,000 held-out cases demonstrate CORE-Acu's superior entity fidelity and reasoning quality. Crucially, CORE-Acu achieved 0/1,000 observed safety violations (95\% CI: 0--0.37\%), whereas GPT-4o exhibited an 8.5\% violation rate under identical rules. These results establish CORE-Acu as a robust neuro-symbolic framework for acupuncture clinical decision support, guaranteeing both reasoning auditability and strict safety compliance.

针灸AI推理可解释知识图谱安全验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。