arXiv:2608.21482eess.IVcs.AI2026-08

用解剖一致性门控提升骨折分类鲁棒性,避免错误信息干扰。

Reliability- and Anatomy-Consistency-Aware Multimodal Learning for Robust Fracture Classification from Bangladeshi Radiographs

论文配图:Reliability- and Anatomy-Consistency-Aware Multimodal Learning for Robust Fracture Classification from Bangladeshi Radiographs
图 1 · 摘自论文原文
  • 引入解剖一致性门控,动态过滤不匹配的临床信息
  • 在五折交叉验证中,宏平均F1达0.6046,优于图像仅模型
  • 适用于临床数据不完整或错配场景,适合医疗影像系统

背景:多模态骨折分类可利用患者和解剖元数据,但当上下文信息缺失或不匹配时可能变得脆弱。方法:基于孟加拉国OrthoFrac-XR数据集的1493张放射片,采用泄漏安全的年龄、性别、骨骼类型和侧别信息。将ConvNeXt图像编码器与临床全连接网络通过拼接、后期融合、可靠性门控残差融合及分层状态-位置建模相结合。额外提出解剖一致性门控,当图像侧向解剖预测与报告骨骼类型冲突时,减弱元数据修正作用。结果:在五折三种子实验中,分层残差融合获得0.6046 ± 0.0279的宏平均F1,优于仅图像学习的0.5727 ± 0.0270;同时Brier得分从0.5239降至0.4948。在五折鲁棒性测试中,解剖一致性融合将元数据乱序下的宏平均F1损失从0.0567降至0.0203,虽干净数据表现略低。无骨骼类型输入时,辅助解剖监督使宏平均F1从0.5620 ± 0.0330提升至0.5899 ± 0.0289。结论:结构化上下文有助于骨折分类,一致性感知门控可限制错误元数据带来的损害。观察到的性能-鲁棒性权衡及缺乏患者级标识符,提示需外部和前瞻性验证。

原文摘要 · Abstract (English)

Background: Multimodal fracture classifiers may benefit from patient and anatomical metadata, but they can also become brittle when contextual information is missing or mismatched. Methods: We studied 1493 radiographs from the Bangladeshi OrthoFrac-XR dataset using leakage-safe age, sex, bone type, and laterality. A ConvNeXt image encoder was combined with a clinical multilayer perceptron through concatenation, late fusion, reliability-gated residual fusion, and a hierarchical state-location formulation. We additionally introduced an anatomy-consistency gate that attenuates metadata corrections when an image-side anatomical prediction disagrees with the reported bone type. Results: Across five folds and three seeds, hierarchical residual fusion achieved a macro-F1 of 0.6046 +/- 0.0279, compared with 0.5727 +/- 0.0270 for image-only learning, while improving the Brier score from 0.5239 to 0.4948. In a five-fold robustness experiment, anatomy-consistency fusion reduced the macro-F1 loss under shuffled metadata from 0.0567 to 0.0203 relative to ordinary residual fusion, although its clean-data macro-F1 was lower. Without bone type at inference, auxiliary anatomy supervision improved macro-F1 from 0.5620 +/- 0.0330 to 0.5899 +/- 0.0289. Conclusions: Structured context improves fracture classification, and consistency-aware gating limits harm from mismatched metadata. The observed clean-performance-robustness trade-off and the absence of patient-level identifiers motivate external and prospective validation.

骨折分类多模态学习医学影像鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。