用多智能体系统自动发现并修复医学影像模型性能下降问题
ReclAIm: A Multi-Agent Framework for Monitoring and Correcting Performance Decline in Medical Imaging AI
- 通过自然语言交互的多智能体架构,自动监控模型表现
- 性能下降最高达40.6%时,微调后恢复至基线水平的98%以内
- 适合医疗AI研究者与临床部署团队使用,降低维护成本
目的:开发并评估一种基于大语言模型的多智能体框架(ReclAIm),用于自动化监测、检测和纠正医学图像分类模型的性能衰退。方法:ReclAIm是一个基于大语言模型的多智能体系统,通过自然语言交互运行。主控智能体协调三个任务专用智能体,执行性能评估,并在检测到显著性能下降时触发微调流程。微调工作流包含数据增强、类别不平衡处理以及参数锚定正则化策略,以缓解灾难性遗忘。该系统在多个影像数据集上进行了基准测试,包括脑部MRI、胸部CT和胸部放射摄影,数据按60%:20%:20%比例划分为模型开发、推理(性能监控)和微调子集。结果:ReclAIm成功在所有数据集上协同完成训练、评估与性能监控。在18个模型中有8个检测到测试集与推理集之间的性能差异,触发微调流程,有效缩小了性能差距。在心影增大数据集(InceptionV3模型)中,性能下降高达40.6%,经微调后,性能指标恢复至基线值的98%以内。结论:ReclAIm提供了一个自动化监控与针对性微调的原型框架,具备自然语言接口,有助于提升科研与潜在临床应用中的可及性。
原文摘要 · Abstract (English)
Purpose: To develop and evaluate a multi-agent framework (ReclAIm) for automated monitoring, detection, and correction of performance decline in medical image classification models. Materials and Methods: ReclAIm is a large language model-based multi-agent system that operates through natural language interaction. A master agent coordinating three task-specific agents performed performance evaluation and triggered fine-tuning when substantial performance declines were detected. The fine-tuning workflow incorporated data augmentation, class imbalance handling, and a parameter-anchoring regularization strategy to limit catastrophic forgetting. The system was benchmarked using multiple imaging datasets, including brain MRI, chest CT, and chest radiography, partitioned into model development, inference (performance monitoring), and fine-tuning subsets (60%:20%:20%). Results: ReclAIm successfully orchestrated training, evaluation, and performance monitoring across all datasets. Performance discrepancies between test and inference data were detected in 8 of 18 models, prompting fine-tuning workflows that reduced performance gaps. In cases with declines of up to 40.6% (cardiomegaly dataset, InceptionV3), fine-tuning restored performance metrics to within 2% of baseline values. Conclusion: ReclAIm provides a prototype framework for automated monitoring and targeted fine-tuning of medical image classification models, with a natural language interface designed to support accessibility in research and potential clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。