让医疗对话分析更准:用知识图谱引导模型,提升细粒度标注一致性。
EPPC-OASIS: Ontology-Aware Adaptation and Structured Inference Refinement for Electronic Patient-Provider Communication Mining in Secure Messages

- 基于医学术语本体增强模型训练,对齐语义空间与代码结构
- 在多个大模型上实现77.13%代码+子代码F1,比基线高2.12点
- 适合医疗文本挖掘、电子病历分析等需结构化标注的场景
安全的患者-医生通信中包含重要临床行为,但人工标注难以规模化。电子患者-医生沟通(EPPC)框架提供编码本体,但自动化提取仍具挑战,因需保持细粒度代码/子代码结构并扎根于原文。我们提出EPPC-OASIS,一种本体感知的适应方法,结合可部署的推理优化流程以提升标注一致性。该方法在监督微调基础上引入Wasserstein对齐目标,使模型表示邻域与本体导出邻域对齐;推理优化则通过验证、自洽性、混合修正及选择或集成策略缓解残余误差。我们在去标识化的安全通信语料库上评估,对比提示、监督微调、偏好驱动及鲁棒性基线,在多个开源语言模型上测试。最佳可部署流水线在所有模型族中达到77.13% Code+Sub-code F1和63.83% Triplet F1,较最强监督微调基线分别提升1.39和2.12 F1点。结果表明,本体感知适配与结构化推理优化可支持可扩展的回顾性EPPC挖掘,但尚需外部验证方可投入实际使用。
原文摘要 · Abstract (English)
Secure patient-provider messages contain clinically important communication behaviors that are difficult to characterize manually at scale. The Electronic Patient-Provider Communication (EPPC) framework provides an ontology for coding these behaviors, but automated extraction remains challenging because predictions must preserve fine-grained code/sub-code structure while grounding annotations in message text. We developed EPPC-OASIS, an ontology-aware adaptation approach for structured EPPC extraction, and combined it with deployable inference-refinement procedures designed to improve the coherence of final annotations. EPPC-OASIS augments supervised fine-tuning with a Wasserstein alignment objective that encourages alignment between model representation neighborhoods and EPPC ontology-derived neighborhoods, while inference refinement uses verification, self-consistency, hybrid correction, and selection or ensembling to address residual prediction errors. We evaluated the framework on a de-identified corpus of secure patient-provider messages against prompting, supervised fine-tuning, preference-based, and robustness-oriented baselines across multiple open-weight language models. Across model families, the best deployable pipeline achieved 77.13% Code+Sub-code F1 and 63.83% Triplet F1, corresponding to modest but consistent absolute gains of +1.39 and +2.12 F1 points over the strongest supervised fine-tuning baseline. These results suggest that ontology-aware adaptation with structured inference refinement can support scalable retrospective EPPC mining, although external validation is needed before operational use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。