提出以活性成分作为药物推荐的新粒度,兼顾精准与安全。
GRAIN: Molecules Are Not the Right Granularity -- Active-Ingredient Modeling for Safe Medication Recommendation
- 用活性成分替代药品编码作为建模粒度,更贴合药物相互作用机制。
- 在MIMIC-IV上提升多标签指标,同时将药物相互作用率从18.75%降至9.48%。
- 引入比例控制器动态调节安全与精度平衡,适合临床用药安全场景。
从电子健康记录中进行药物推荐需在预测准确性和多重用药带来的药物-药物相互作用(DDI)风险之间取得平衡。现有安全感知推荐系统通常以药品代码或分子结构为粒度,前者将每种药物视为不可分割的单元,后者则细于实际药理交互知识组织方式。本文认为活性成分才是缺失的合适粒度,提出基于活性成分的推荐框架GRAIN。GRAIN采用选择性状态空间主干网络,线性时间处理长期不规则的就诊序列。在此基础上,构建联合目标函数,统一整合三类知识:药品级DDI图、通过RxNorm映射到活性成分的成分级DDI图,以及来自EHR的共处方图。比例控制器根据验证集中的实际DDI率动态调节准确率与安全性权衡,而非预先固定。在完全匹配设置下(预处理、队列、词汇表、划分和评估代码一致),GRAIN在MIMIC-IV上超越重实现的MambaHealth基线,在所有标准多标签指标上均取得提升(Jaccard从0.4488升至0.4983,PRAUC从0.6911升至0.7485,F1从0.5989升至0.6453),同时将药品级DDI率由0.1875降至0.0948。进一步定义成分级DDI率这一原方法无法观测的安全指标。结果表明,成分级归一化恢复了被代码级聚合掩盖的预测信号,且与精准序列建模互补而非竞争。
原文摘要 · Abstract (English)
Medication recommendation from electronic health records must balance predictive accuracy against the risk of adverse drug-drug interactions (DDIs) under polypharmacy. Existing safety-aware recommenders operate at one of two granularities: the drug code, which treats each medication as an indivisible token, or the molecular substructure, which is finer than pharmacological interaction knowledge is actually organized. We argue that the active ingredient is the missing granularity, and introduce GRAIN, a medication recommendation framework built around it. GRAIN encodes longitudinal patient trajectories (diagnoses, procedures, past medications) with a selective state space backbone that handles long, irregular visit sequences in linear time. On top of it we introduce a joint objective unifying three knowledge sources aligned to a common medication vocabulary: a drug-level DDI graph, an ingredient-level DDI graph obtained by normalizing medication codes to active ingredients via RxNorm, and an EHR-derived co-prescription graph. A proportional controller adapts the accuracy-safety trade-off to the observed validation DDI rate rather than fixing it a priori. Under strictly matched settings -- identical preprocessing, cohort, vocabulary, split, and evaluation code -- GRAIN improves over a re-implemented MambaHealth baseline on MIMIC-IV across all standard multi-label metrics (Jaccard 0.4488 to 0.4983, PRAUC 0.6911 to 0.7485, F1 0.5989 to 0.6453) while reducing the drug-level DDI rate from 0.1875 to 0.0948. We further define an ingredient-level DDI rate, a safety measure invisible to drug-code-level evaluation. The results indicate that ingredient-level normalization recovers predictive signal erased by code-level aggregation, and that it is complementary to, rather than in competition with, accurate sequence modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。