arXiv:2410.19643cs.LGcs.AI2024-10被引 5

解决医学数据跨中心差异中的标签泄漏问题,提升模型泛化能力。

Impact of Leakage on Data Harmonization in Machine Learning Pipelines in Class Imbalance Across Sites

  • 用伪标签模拟目标,避免数据泄露。
  • 在真实MRI数据上表现与传统方法相当。
  • 特别适合标签分布不均的多中心研究。

机器学习模型依赖大规模数据,但生物医学数据收集成本高,常需合并多来源数据。然而不同来源数据存在非期望的站点特异性差异。现有基于ComBat的方法虽用于消除站点差异,但在各站点类别不平衡时易产生数据泄露。本研究发现此类方法存在严重泄露问题,并提出新方法PrettYharmonize,通过假装目标标签来实现数据调和。我们使用受控数据集评估调和效果,并在真实世界MRI与临床数据上对比验证。结果表明,PrettYharmonize在保持性能的同时有效避免了数据泄露,尤其在站点与目标相关性高的场景中表现更优。

原文摘要 · Abstract (English)

Machine learning (ML) models benefit from large datasets. Collecting data in biomedical domains is costly and challenging, hence, combining datasets has become a common practice. However, datasets obtained under different conditions could present undesired site-specific variability. Data harmonization methods aim to remove site-specific variance while retaining biologically relevant information. This study evaluates the effectiveness of popularly used ComBat-based methods for harmonizing data in scenarios where the class balance is not equal across sites. We find that these methods struggle with data leakage issues. To overcome this problem, we propose a novel approach PrettYharmonize, designed to harmonize data by pretending the target labels. We validate our approach using controlled datasets designed to benchmark the utility of harmonization. Finally, using real-world MRI and clinical data, we compare leakage-prone methods with PrettYharmonize and show that it achieves comparable performance while avoiding data leakage, particularly in site-target-dependence scenarios.

数据调和跨中心学习类别不平衡医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。