arXiv:2411.01653cs.LGcs.AI2024-11

用训练动态检测医疗数据质量,发现现有方法不适用。

Diagnosing Medical Datasets with Training Dynamics

  • 通过训练过程中的学习行为分类数据难易与模糊性
  • 医疗问答数据中难学与模糊样本占比高,影响模型性能
  • 该方法在医学领域效果不佳,不具通用性

本研究探索利用训练动态作为自动化手段,替代人工标注来评估训练数据质量。采用的框架是Data Maps,可将数据点分类为易学、难学和模糊三类(Swayamdipta et al., 2020)。已有研究表明,难学样本常含错误,模糊样本显著影响模型训练。为验证该结论的可靠性,我们基于一个具有挑战性的医疗问答数据集复现实验,该任务不仅要求文本理解,还需掌握详尽医学知识,进一步提升难度。全面评估表明,Data Maps框架在医疗领域难以应对特定挑战,不具备可行性与可迁移性。

原文摘要 · Abstract (English)

This study explores the potential of using training dynamics as an automated alternative to human annotation for evaluating the quality of training data. The framework used is Data Maps, which classifies data points into categories such as easy-to-learn, hard-to-learn, and ambiguous (Swayamdipta et al., 2020). Swayamdipta et al. (2020) highlight that difficult-to-learn examples often contain errors, and ambiguous cases significantly impact model training. To confirm the reliability of these findings, we replicated the experiments using a challenging dataset, with a focus on medical question answering. In addition to text comprehension, this field requires the acquisition of detailed medical knowledge, which further complicates the task. A comprehensive evaluation was conducted to assess the feasibility and transferability of the Data Maps framework to the medical domain. The evaluation indicates that the framework is unsuitable for addressing datasets' unique challenges in answering medical questions.

数据质量医疗AI训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。