arXiv:2510.00024cs.SIcs.AI2025-10被引 6

用AI自动完成流行病建模研究全流程,生成完整论文。

EpidemIQs: Prompt-to-Paper LLM Agents for Epidemic Modeling and Analysis

  • 分五阶段多智能体协作,科学家与任务专家分工执行
  • 平均任务成功率79%,单次研究成本约1.57美元
  • 适合科研人员快速生成建模报告,提升研究效率

大型语言模型(LLMs)为加速复杂跨学科研究提供了新机遇。流行病建模涉及网络科学、动力系统、流行病学与随机模拟,是推动LLM自动化的重要领域。我们提出EpidemIQs,一种新型多智能体LLM框架,通过五个预设研究阶段,整合用户输入,自主完成文献综述、分析推导、网络建模、机制建模、随机模拟、数据可视化与分析,并生成结构化论文报告。框架包含两类智能体:科学家智能体负责规划、协调、反思与最终结果生成;任务专家智能体专注单一任务,作为工具支持科学家智能体。使用GPT-4.1和GPT-4.1 Mini分别作为科学家与任务专家的骨干模型,平均总令牌消耗870K,单次研究成本约1.57美元,成功完成全部阶段并生成完整报告。我们在多个流行病情景下评估了EpidemIQs,衡量计算成本、工作流可靠性、任务成功率及LLM与人工专家评审,结果显示平均任务成功率为79%。与采用相同提示与工具的迭代单智能体方法对比,EpidemIQs表现更优,验证了其高效性与稳定性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) offer new opportunities to accelerate complex interdisciplinary research domains. Epidemic modeling, characterized by its complexity and reliance on network science, dynamical systems, epidemiology, and stochastic simulations, represents a prime candidate for leveraging LLM-driven automation. We introduce EpidemIQs, a novel multi-agent LLM framework that integrates user inputs and autonomously conducts literature review, analytical derivation, network modeling, mechanistic modeling, stochastic simulations, data visualization and analysis, and finally documentation of findings in a structured manuscript, through five predefined research phases. We introduce two types of agents: a scientist agent for planning, coordination, reflection, and generation of final results, and a task-expert agent to focus exclusively on one specific duty serving as a tool to the scientist agent. The framework consistently generated complete reports in scientific article format. Specifically, using GPT 4.1 and GPT 4.1 Mini as backbone LLMs for scientist and task-expert agents, respectively, the autonomous process completes with average total token usage 870K at a cost of about $1.57 per study, successfully executing all phases and final report. We evaluate EpidemIQs across several different epidemic scenarios, measuring computational cost, workflow reliability, task success rate, and LLM-as-Judge and human expert reviews to estimate the overall quality and technical correctness of the generated results. Through our experiments, the framework consistently addresses evaluation scenarios with an average task success rate of 79%. We compare EpidemIQs to an iterative single-agent LLM, benefiting from the same system prompts and tools, iteratively planning, invoking tools, and revising outputs until task completion. The comparisons suggest a consistently higher performance of EpidemIQs.

流行病建模多智能体自动化科研LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。