研究医生选迁移数据的直觉,发现相似性不等于效果好。
Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification
- 通过调查了解研究人员如何选医学图像迁移数据
- 相似度高未必性能好,实际表现与直觉不符
- 适合关注医疗AI落地的从业者和研究者
迁移学习在医学影像中至关重要,但源数据集的选择常依赖研究者的直觉而非系统原则,可能影响算法泛化能力及患者结果。本研究通过任务型问卷调查机器学习从业者,从人机交互角度分析其源数据选择行为。结果显示,选择受任务特性、社区惯例、数据属性及感知视觉/语义相似性影响;但相似性评分与预期性能并不总一致,挑战了“越相似越好”的传统观念。此外,伦理与公平性考量在源数据描述中几乎缺失。参与者常使用模糊术语,表明亟需更清晰的定义与工具支持。本文通过厘清这些启发式方法,并提出迁移学习因素的概念框架,为更系统的源数据选择提供实践指导。
原文摘要 · Abstract (English)
Transfer learning is crucial for medical imaging, yet the selection of source datasets often relies on researchers' intuition rather than systematic principles, which can impact the generalizability of algorithms and, thus, patient outcomes. This study investigates these decisions through a task-based survey with machine learning practitioners. Unlike prior work that benchmarks models and experimental setups, we take a human-computer interaction (HCI) perspective on how practitioners select source datasets. Our findings indicate that choices are task-dependent and influenced by community practices, dataset properties, and computational (data embedding), or perceived visual or semantic similarity. However, similarity ratings and expected performance are not always aligned, challenging a traditional "more similar is better" view. Moreover, ethical and fairness considerations remain largely absent from source dataset sections. Participants often used ambiguous terminology, which suggests a need for clearer definitions and tools to make them explicit and usable. By clarifying these heuristics and introducing a conceptual framework of transfer learning factors, this work provides practical insights for more systematic source selection in transfer learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。