用多任务模型自动评估医疗器械召回严重程度与根本原因。
RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage

- 融合文本与结构化数据的双通道模型,同时预测召回等级和原因类别。
- 多任务模型性能显著优于单任务基线,严重程度分类准确率达96.3%。
- 适合监管机构、医疗安全研究者用于大规模召回风险智能筛查。
医疗器械召回是保障患者安全的关键监管手段。随着FDA召回记录数量激增,事后召回分诊、严重程度评估与根本原因分析面临挑战。现有研究多单独处理召回发生预测或原因分析,联合建模严重程度与原因类别的工作较少。本文基于54,165条来自openFDA的医疗器械召回记录(2002年至2025年10月),评估了经典机器学习与基于提升的模型在严重程度和原因类别预测中的表现。随后提出RecallRisk-BERT,一种多任务模型,结合PubMedBERT生成的召回描述文本表示与产品编码、法规编号、医学专业等结构化特征的嵌入表示,同步预测召回严重程度(I/II/III类)与整合后的9类根本原因。使用准确率、宏平均精确率、召回率、F1值及ROC-AUC进行评估。单任务下,基于LightGBM的文本-表格配置表现最佳,准确率为0.963,宏-F1为0.856,ROC-AUC为0.974。多任务设置中,RecallRisk-BERT显著超越单任务PubMedBERT基线。模型推导的风险排序与实际原因严重性模式高度一致(rho = 0.983, p = 1.936e-6)。结果表明,文本-表格学习可支持可扩展的召回后分诊、监管决策支持与基于模型的根本原因风险分析。
原文摘要 · Abstract (English)
Medical device recalls are a critical regulatory mechanism for protecting patient safety. The growing volume of FDA recall records presents challenges in post-report recall triage, severity assessment, and root-cause interpretation. Existing studies mostly address recall occurrence prediction or root-cause analysis separately, while joint modeling of recall severity and root-cause categories has received limited attention. We develop an automated recall triage framework using 54,165 FDA medical device recall records from openFDA, covering the period from 2002 to October 2025. We first evaluate classical machine learning and boosting-based models for recall severity and root-cause category prediction. We then develop RecallRisk-BERT, a multi-task model that combines PubMedBERT-based textual representations of recall narratives with embedding-based representations of structured categorical features, including product code, regulation number, and medical specialty. The model simultaneously predicts recall severity (Class I/II/III) and a consolidated root-cause category (9 classes). Performance was evaluated using accuracy, macro-averaged precision, recall, F1-score, and ROC-AUC. In single-task severity prediction, our LightGBM-based text--tabular configuration achieved the strongest performance, with an accuracy of 0.963, macro-F1 of 0.856, and ROC-AUC of 0.974. In the multi-task setting, RecallRisk-BERT substantially outperformed the single-task PubMedBERT baseline. Model-derived risk rankings were strongly consistent with observed root-cause severity patterns (rho = 0.983, p = 1.936e-6). These findings indicate that text--tabular learning can support scalable post-report recall triage, regulatory decision support, and model-based root-cause risk analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。