用合成数据暴露传统多实例学习在关联信息利用上的缺陷
Synthetic Data Reveals Generalization Gaps in Correlated Multiple Instance Learning
- 设计合成任务,强调相邻实例间关系的重要性
- 对比最优贝叶斯估计器,发现现有方法性能仍有差距
- 即使使用1万样本,新方法仍无法达到理论最优
多实例学习(MIL)常用于医学影像分类,通过处理图像块或切片来判断高分辨率2D图像或3D体数据。然而,传统MIL方法将实例独立处理,忽略了邻近图像块或切片之间的上下文关系,而这些关系在实际应用中至关重要。本文设计了一个合成分类任务,其中利用相邻实例特征是准确预测的关键。通过与可闭式求解的最优贝叶斯估计器对比,量化评估了现成MIL方法的性能局限。实验表明,即使在每个样本包含大量实例、训练样本达一万的情况下,较新的相关性MIL方法仍未能达到最佳可能性能。
原文摘要 · Abstract (English)
Multiple instance learning (MIL) is often used in medical imaging to classify high-resolution 2D images by processing patches or classify 3D volumes by processing slices. However, conventional MIL approaches treat instances separately, ignoring contextual relationships such as the appearance of nearby patches or slices that can be essential in real applications. We design a synthetic classification task where accounting for adjacent instance features is crucial for accurate prediction. We demonstrate the limitations of off-the-shelf MIL approaches by quantifying their performance compared to the optimal Bayes estimator for this task, which is available in closed-form. We empirically show that newer correlated MIL methods still do not achieve the best possible performance when trained with ten thousand training samples, each containing many instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。