arXiv:2604.16362cs.LGcs.AI2026-04

SetFlow通过生成表示集提升弱监督医学图像分类效果

SetFlow: Generating Structured Sets of Representations for Multiple Instance Learning

论文配图:SetFlow: Generating Structured Sets of Representations for Multiple Instance Learning
图 1 · 摘自论文原文
  • 直接在表示空间建模整个数据包,捕捉袋内实例关联
  • 生成样本逼近真实分布,增强下游任务性能1.8%以上
  • 适合数据稀缺或隐私敏感场景的合成数据生成

数据稀缺与弱监督限制了机器学习在真实场景中的表现,如乳腺钼靶影像中常采用多实例学习(MIL)框架。尽管当前基础模型提供强大的语义表示,但现有方法仅在实例层面进行增强,难以捕捉袋内依赖关系。本文提出SetFlow,一种直接在表示空间建模整个MIL数据包(即集合)的生成架构。该方法结合流匹配范式与受Set Transformer启发的设计,可处理排列不变输入并捕捉袋内实例间交互。模型同时以类别标签和输入尺度为条件,生成语义一致且连贯的表示集合。在大规模乳腺钼靶基准上,使用最先进的MIL-PF分类流水线评估,生成样本与原始数据分布高度一致,且作为增强数据能提升下游性能。仅用合成数据训练亦取得有竞争力结果,证明表示空间生成建模对数据稀缺及隐私敏感任务的有效性。

原文摘要 · Abstract (English)

Data scarcity and weak supervision continue to limit the performance of machine learning models in many real-world applications, such as mammography, where Multiple Instance Learning (MIL) often offers the best formulation. While recent foundation models provide strong semantic representations out of the box, effective augmentation of such representations of MIL data remains limited, as existing methods operate at the instance level and fail to capture intra-bag dependencies. In this work, we introduce SetFlow, a generative architecture that models entire MIL bags (i.e., sets) directly in the representation space. Our approach leverages the flow matching paradigm combined with a Set Transformer-inspired design, enabling it to handle permutation-invariant inputs while capturing interactions between instances within each bag. The model is conditioned on both class labels and input scale, allowing it to generate coherent and semantically consistent sets of representations. We evaluate SetFlow on a large-scale mammography benchmark using a state-of-the-art MIL-PF classification pipeline. The generated samples are shown to closely match the original data distribution and even improve downstream performance when used for augmentation. Furthermore, training on synthetic data alone shows competitive results, demonstrating the effectiveness of representation-space generative modeling for data-scarce and privacy-sensitive tasks.

多实例学习生成模型医学影像数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。