arXiv:2511.01140stat.MLcs.AI2025-11被引 1

为少样本多模态医学影像提供理论框架,解释模型为何有效且如何改进诊断可靠性。

Few-Shot Multimodal Medical Imaging: A Theoretical Framework

  • 基于概率学习理论,推导出可靠性能所需的最少标注样本数。
  • 证明多模态信息可降低模型复杂度,减少预测不确定性并提升解释稳定性。
  • 适合关注数据效率、可信诊断与模型可解释性的医疗AI研究者。

医学影像常面临标注数据稀缺问题,尤其在罕见病和资源匮乏的临床环境中。现有少样本与多模态方法虽能提升性能,但缺乏理论解释。本文提出统一的少样本多模态医学影像理论框架,联合刻画样本复杂度、不确定性量化与可解释性。利用PAC学习、VC理论与PAC贝叶斯分析,推导出保证可靠性能所需的最小标注样本量,并证明互补模态通过信息增益项降低有效容量。进一步提出解释稳定性形式化度量,证明解释方差随样本数倒数速率下降。还构建链式思维的序贯贝叶斯解释模型,展示后验逐步收缩过程。通过可控多模态数据集与加性CNN-MLP融合模型验证,结果确认了预测收益、大样本下模态干扰现象及预测不确定性持续缩小。该框架为低资源环境下高效、可靠、可解释的诊断模型设计提供了理论基础。

原文摘要 · Abstract (English)

Medical imaging often operates under limited labeled data, especially in rare disease and low resource clinical environments. Existing multimodal and meta learning approaches improve performance in these settings but lack a theoretical explanation of why or when they succeed. This paper presents a unified theoretical framework for few shot multimodal medical imaging that jointly characterizes sample complexity, uncertainty quantification, and interpretability. Using PAC learning, VC theory, and PAC Bayesian analysis, we derive bounds that describe the minimum number of labeled samples required for reliable performance and show how complementary modalities reduce effective capacity through an information gain term. We further introduce a formal metric for explanation stability, proving that explanation variance decreases at an inverse n rate. A sequential Bayesian interpretation of Chain of Thought reasoning is also developed to show stepwise posterior contraction. To illustrate these ideas, we implement a controlled multimodal dataset and evaluate an additive CNN MLP fusion model under few shot regimes, confirming predicted multimodal gains, modality interference at larger sample sizes, and shrinking predictive uncertainty. Together, the framework provides a principled foundation for designing data efficient, uncertainty aware, and interpretable diagnostic models in low resource settings.

少样本学习多模态医学影像可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。