用智能代理自动完成医疗联邦学习全流程,减少人工干预。
FedAgentBench: Towards Automating Real-world Federated Medical Image Analysis with Server-Client LLM Agents
- 构建客户端与服务器端大模型代理协同的自动化联邦学习框架。
- 涵盖40种算法、201个数据集,覆盖6类医学影像模态。
- 验证大模型在复杂任务中自主决策能力,揭示当前局限性。
联邦学习(FL)可在不共享敏感患者数据的前提下实现多医疗机构协作建模。然而,实际部署常受制于大量人为操作挑战:包括选择合适医院客户端、协调服务器与客户端、客户端数据预处理、跨机构数据与标签标准化,以及根据用户指令和跨客户端数据特征选择合适的联邦算法。现有研究普遍忽视这些实际调度难题。为此,我们提出一个代理驱动的联邦学习框架,覆盖从客户端筛选到训练完成的全流程,并构建名为FedAgentBench的基准,评估大语言模型(LLM)代理在医疗联邦学习中自主协调的能力。框架集成40种联邦算法,适配不同任务需求与跨机构特性;引入201个精心设计的数据集,模拟6类真实医疗环境:皮肤镜、超声、眼底、病理、MRI和X光。我们评估了14个开源与10个专有大模型(涵盖小、中、大尺度)的代理性能。结果显示,尽管如GPT-4.1和DeepSeek V3等强模型能自动化部分流程,但基于隐含目标的复杂、相互依赖的任务仍对最先进模型构成挑战。
原文摘要 · Abstract (English)
Federated learning (FL) allows collaborative model training across healthcare sites without sharing sensitive patient data. However, real-world FL deployment is often hindered by complex operational challenges that demand substantial human efforts. This includes: (a) selecting appropriate clients (hospitals), (b) coordinating between the central server and clients, (c) client-level data pre-processing, (d) harmonizing non-standardized data and labels across clients, and (e) selecting FL algorithms based on user instructions and cross-client data characteristics. However, the existing FL works overlook these practical orchestration challenges. These operational bottlenecks motivate the need for autonomous, agent-driven FL systems, where intelligent agents at each hospital client and the central server agent collaboratively manage FL setup and model training with minimal human intervention. To this end, we first introduce an agent-driven FL framework that captures key phases of real-world FL workflows from client selection to training completion and a benchmark dubbed FedAgentBench that evaluates the ability of LLM agents to autonomously coordinate healthcare FL. Our framework incorporates 40 FL algorithms, each tailored to address diverse task-specific requirements and cross-client characteristics. Furthermore, we introduce a diverse set of complex tasks across 201 carefully curated datasets, simulating 6 modality-specific real-world healthcare environments, viz., Dermatoscopy, Ultrasound, Fundus, Histopathology, MRI, and X-Ray. We assess the agentic performance of 14 open-source and 10 proprietary LLMs spanning small, medium, and large model scales. While some agent cores such as GPT-4.1 and DeepSeek V3 can automate various stages of the FL pipeline, our results reveal that more complex, interdependent tasks based on implicit goals remain challenging for even the strongest models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。