用闭环数据引擎让医学影像评估模型越改越好,只用1万标注就接近人类水平。
MedQ-Engine: A Closed-Loop Data Engine for Evolving MLLMs in Medical Image Quality Assessment
- 通过聚类发现模型失败案例,用这些案例检索图像并引导人工标注。
- 80亿参数模型经训练后超越GPT-4o超13%,与人类差距仅4.34%。
- 自适应迭代优化,标注效率是随机采样的4倍以上,适合医疗AI持续改进。
医学影像质量评估(Med-IQA)是临床AI部署的前提,但多模态大语言模型仍远未达到人类专家水平,尤其在需要结合临床推理的描述性评估方面。现有方法受限于高成本的描述性标注及一次性数据采集无法适配模型演进的弱点。为此,我们提出MedQ-Engine,一个闭环数据引擎:通过数据驱动聚类识别模型失败原型,以原型为锚点在百万级图像库中检索,并结合渐进式人机协同标注;再通过质量保障的微调实现模型演化,形成自我提升循环。模型在互补的感知与描述任务上评估。熵引导路由机制有效分配标注资源,降低标注成本。在五种医学成像模态上的实验表明,该方法使一个80亿参数模型超越GPT-4o超过13%,与人类专家差距缩小至4.34%,仅使用10,000个标注样本,样本效率高于随机采样4倍以上。
原文摘要 · Abstract (English)
Medical image quality assessment (Med-IQA) is a prerequisite for clinical AI deployment, yet multimodal large language models (MLLMs) still fall substantially short of human experts, particularly when required to provide descriptive assessments with clinical reasoning beyond simple quality scores. However, improving them is hindered by the high cost of acquiring descriptive annotations and by the inability of one-time data collection to adapt to the model's evolving weaknesses. To address these challenges, we propose MedQ-Engine, a closed-loop data engine that iteratively evaluates the model to discover failure prototypes via data-driven clustering, explores a million-scale image pool using these prototypes as retrieval anchors with progressive human-in-the-loop annotation, and evolves through quality-assured fine-tuning, forming a self-improving cycle. Models are evaluated on complementary perception and description tasks. An entropy-guided routing mechanism triages annotations to minimize labeling cost. Experiments across five medical imaging modalities show that MedQ-Engine elevates an 8B-parameter model to surpass GPT-4o by over 13% and narrow the gap with human experts to only 4.34%, using only 10K annotations with more than 4x sample efficiency over random sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。