arXiv:2602.07154cs.LGcs.AI2026-02被引 3

提出匹配方法缓解数据异质性,提升零样本医学异常检测性能

Beyond Pooling: Matching for Robust Generalization under Data Heterogeneity

  • 通过自适应中心点匹配样本,动态优化表示分布
  • 在非高斯、多模态真实场景中优于传统池化与随机采样
  • 特别适合极端异质性的零样本医疗异常检测任务

跨领域异质数据池化是表征学习中的常见策略,但简单拼接会放大分布不对称性,导致偏差估计,尤其在需要零样本泛化时。本文提出一种匹配框架,基于自适应中心点选择样本并迭代优化表示分布。该方法兼具双重稳健性,并引入倾向得分匹配以排除混杂域(异质性主因),使匹配比朴素池化或均匀采样更鲁棒。理论与实证分析表明,相较于朴素池化与均匀采样,匹配在不对称元分布下表现更优,且可扩展至非高斯与多模态真实场景。最重要的是,这些改进在零样本医学异常检测中得以体现,这是数据异质性与不对称性的极端形式。代码已开源:https://github.com/AyushRoy2001/Beyond-Pooling。

原文摘要 · Abstract (English)

Pooling heterogeneous datasets across domains is a common strategy in representation learning, but naive pooling can amplify distributional asymmetries and yield biased estimators, especially in settings where zero-shot generalization is required. We propose a matching framework that selects samples relative to an adaptive centroid and iteratively refines the representation distribution. The double robustness and the propensity score matching for the inclusion of data domains make matching more robust than naive pooling and uniform subsampling by filtering out the confounding domains (the main cause of heterogeneity). Theoretical and empirical analyses show that, unlike naive pooling or uniform subsampling, matching achieves better results under asymmetric meta-distributions, which are also extended to non-Gaussian and multimodal real-world settings. Most importantly, we show that these improvements translate to zero-shot medical anomaly detection, one of the extreme forms of data heterogeneity and asymmetry. The code is available on https://github.com/AyushRoy2001/Beyond-Pooling.

数据异质性零样本学习医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。