构建医疗协作模拟平台,评测大模型群体智能演化能力
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence
- 以医生与患者交互方式让大模型群体智能自我进化
- 实验证明可有效提升医疗专业能力和系统效率
- 适合研究群体智能、AI医疗评估的学者与开发者
基于大语言模型的群体智能(LLM-based Collective Intelligence, CI)为突破数据瓶颈、持续提升大模型代理能力提供了新路径。然而,当前缺乏专门用于演化与评测此类群体智能的开放环境。为此,我们提出OpenHospital,一个互动式实验场,医生代理通过与患者代理交互实现群体智能的持续演化。该平台采用‘数据内置于代理自身’范式,能快速增强代理能力,并提供针对医学专业水平与系统效率的稳健评估指标。实验表明,OpenHospital在促进和量化群体智能方面均具有效性。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based Collective Intelligence (CI) presents a promising approach to overcoming the data wall and continuously boosting the capabilities of LLM agents. However, there is currently no dedicated arena for evolving and benchmarking LLM-based CI. To address this gap, we introduce OpenHospital, an interactive arena where physician agents can evolve CI through interactions with patient agents. This arena employs a data-in-agent-self paradigm that rapidly enhances agent capabilities and provides robust evaluation metrics for benchmarking both medical proficiency and system efficiency. Experiments demonstrate the effectiveness of OpenHospital in both fostering and quantifying CI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。