用合成图像预训练AI,提升犬类骨骼疾病诊断准确率。
Enhancing Canine Musculoskeletal Diagnoses: Leveraging Synthetic Image Data for Pre-Training AI-Models on Visual Documentations
- 用合成图像模拟真实病历影像,解决数据不足问题。
- 小样本下诊断准确率提升约10%,大样本效果不明显。
- 适合标注数据稀缺的医疗AI开发场景。
犬类骨骼系统检查在兽医实践中极具挑战性。本文提出一种新方法,通过可视化方式高效记录犬只状况,但因该可视化文档为全新形式,缺乏训练数据。为此,研究探索了使用模仿真实病历影像的合成数据来预训练AI模型的潜力。首先构建包含3个类别的基础数据集,随后生成含36个类别的更复杂数据集,并用于预训练模型。评估阶段创建了250份人工制作的可视化文档数据集,涵盖五种疾病,另取25例子集进行测试。结果表明,在仅25例的小样本评估中,使用合成图像预训练使诊断准确率提升约10%;但在250例的大样本中未见显著优势,说明合成数据预训练的优势主要体现在样本极少时。本研究为缓解数据匮乏带来的限制提供了有效策略,其方法可推广至其他医疗影像诊断领域。
原文摘要 · Abstract (English)
The examination of the musculoskeletal system in dogs is a challenging task in veterinary practice. In this work, a novel method has been developed that enables efficient documentation of a dog's condition through a visual representation. However, since the visual documentation is new, there is no existing training data. The objective of this work is therefore to mitigate the impact of data scarcity in order to develop an AI-based diagnostic support system. To this end, the potential of synthetic data that mimics realistic visual documentations of diseases for pre-training AI models is investigated. We propose a method for generating synthetic image data that mimics realistic visual documentations. Initially, a basic dataset containing three distinct classes is generated, followed by the creation of a more sophisticated dataset containing 36 different classes. Both datasets are used for the pre-training of an AI model. Subsequently, an evaluation dataset is created, consisting of 250 manually created visual documentations for five different diseases. This dataset, along with a subset containing 25 examples. The obtained results on the evaluation dataset containing 25 examples demonstrate a significant enhancement of approximately 10% in diagnosis accuracy when utilizing generated synthetic images that mimic real-world visual documentations. However, these results do not hold true for the larger evaluation dataset containing 250 examples, indicating that the advantages of using synthetic data for pre-training an AI model emerge primarily when dealing with few examples of visual documentations for a given disease. Overall, this work provides valuable insights into mitigating the limitations imposed by limited training data through the strategic use of generated synthetic data, presenting an approach applicable beyond the canine musculoskeletal assessment domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。