用AI自动生成医学影像标注,让医生只改错,效率提升十倍。
Label Critic: Design Data Before Models
- 通过自动对比多个AI标注,选出最优标签,减少人工工作量。
- 无需微调大模型,96.5%准确率选出最佳标注,单次比对仅需15秒。
- 适合需要高效构建高质量医学数据集的团队或研究者。
随着医疗数据集快速扩张,详细标注人体结构变得愈发昂贵且耗时。我们发现,让放射科医生从头标注并不必要,现有AI模型可大幅自动化该过程。遵循‘不用大锤敲小钉’的原则,我们提出只需医生审查并修正最优AI标签中的错误即可。为此开发了名为Label Critic的工具,通过持续成对比较评估标签质量。实验表明,结合图像-提示对,预训练的大视觉语言模型(LVLM)在无额外微调情况下,96.5%准确率完成配对比较。将人工标注时间(每扫描30–60分钟)转为自动对比任务(每扫描15秒),使医生工作量降低一个数量级。当最优AI标签足够准确(81%,依解剖结构而定)时,直接作为数据集金标准,低质标签自动剔除。此外,当无备选标签时,该工具仍能以71.8%准确率评估单个标签质量,若估计质量低(19%依赖结构),则提示医生介入修改。
原文摘要 · Abstract (English)
As medical datasets rapidly expand, creating detailed annotations of different body structures becomes increasingly expensive and time-consuming. We consider that requesting radiologists to create detailed annotations is unnecessarily burdensome and that pre-existing AI models can largely automate this process. Following the spirit don't use a sledgehammer on a nut, we find that, rather than creating annotations from scratch, radiologists only have to review and edit errors if the Best-AI Labels have mistakes. To obtain the Best-AI Labels among multiple AI Labels, we developed an automatic tool, called Label Critic, that can assess label quality through tireless pairwise comparisons. Extensive experiments demonstrate that, when incorporated with our developed Image-Prompt pairs, pre-existing Large Vision-Language Models (LVLM), trained on natural images and texts, achieve 96.5% accuracy when choosing the best label in a pair-wise comparison, without extra fine-tuning. By transforming the manual annotation task (30-60 min/scan) into an automatic comparison task (15 sec/scan), we effectively reduce the manual efforts required from radiologists by an order of magnitude. When the Best-AI Labels are sufficiently accurate (81% depending on body structures), they will be directly adopted as the gold-standard annotations for the dataset, with lower-quality AI Labels automatically discarded. Label Critic can also check the label quality of a single AI Label with 71.8% accuracy when no alternatives are available for comparison, prompting radiologists to review and edit if the estimated quality is low (19% depending on body structures).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。