用图文融合模型自动分类医疗设备风险等级,提升监管效率与安全性。
Toward Automated Regulatory Decision-Making: Trustworthy Medical Device Risk Classification with Multimodal Transformers and Self-Training
- 结合文本与图像的跨模态注意力机制,捕捉多源信息关联。
- 在真实数据集上达90.4%准确率和97.9% AUROC,显著优于单模态基线。
- 自训练策略有效缓解标注不足问题,适合医疗监管等小样本场景。
准确划分医疗设备风险等级对监管监督和临床安全至关重要。本文提出一种基于Transformer的多模态框架,融合文本描述与视觉信息以预测设备监管分类。模型采用交叉注意力机制捕获模态间依赖关系,并引入自训练策略,在有限监督下提升泛化能力。在真实监管数据集上的实验表明,该方法最高可达90.4%准确率和97.9% AUROC,显著优于仅文本(77.2%)和仅图像(54.8%)的基线。相较于标准多模态融合,自训练使SVM准确率提升3.3个百分点(从87.1%到90.4%),宏平均F1提升1.4点,表明伪标签能有效增强弱监督下的泛化性能。消融实验证实了跨模态注意力与自训练的互补优势。
原文摘要 · Abstract (English)
Accurate classification of medical device risk levels is essential for regulatory oversight and clinical safety. We present a Transformer-based multimodal framework that integrates textual descriptions and visual information to predict device regulatory classification. The model incorporates a cross-attention mechanism to capture intermodal dependencies and employs a self-training strategy for improved generalization under limited supervision. Experiments on a real-world regulatory dataset demonstrate that our approach achieves up to 90.4% accuracy and 97.9% AUROC, significantly outperforming text-only (77.2%) and image-only (54.8%) baselines. Compared to standard multimodal fusion, the self-training mechanism improved SVM performance by 3.3 percentage points in accuracy (from 87.1% to 90.4%) and 1.4 points in macro-F1, suggesting that pseudo-labeling can effectively enhance generalization under limited supervision. Ablation studies further confirm the complementary benefits of both cross-modal attention and self-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。