AI助手自动生成模型,用动态知识库提升科研效率
AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models

- 构建分层外部知识库,动态调用相关领域专家经验
- 在医疗预测竞赛中超越93.4%人类团队,医学基准测试获70.7%奖牌率
- 适合无AI工程背景的科研人员快速搭建高性能模型
AI模型支撑从图像文本处理到生物、物理、化学等科学发现的数据驱动应用。然而,模型开发仍高度依赖人工,需设计架构、构建训练流程并反复优化,对缺乏专业AI工程能力的自然科学家而言难度大。为减轻负担并拓展科学发现中的AI应用,已有自动建模代理提出。但其性能受限于底层大语言模型的静态、过时且稀疏的参数化知识。为此,我们提出AIBuildAI-2,一个具备外部可演进知识系统的增强型代理,用于自动构建AI模型。该知识系统分层组织,将整理后的AI开发知识划分为主题类别下的高层指令与低层文档,代理仅动态加载当前任务相关的上下文,确保每个设计与实现决策基于可验证的外部专长。知识系统初始通过网络收集清洗的AI开发文档建立,并持续通过代理自身完成任务的经验提炼结构化收获,写回知识库。AIBuildAI-2在MLE-Bench上取得70.7%奖牌率,排名第一;在心脏病预测竞赛中,表现优于4,370名人类专家团队中的前6.6%。
原文摘要 · Abstract (English)
AI models underpin data-centric applications from image and text processing to scientific discovery in biology, physics, and chemistry. Yet developing them remains heavily manual, requiring practitioners to design architectures, build training pipelines, and iteratively refine solutions, making it challenging for natural scientists without specialized AI engineering expertise to build the high-performing models their research demands. To reduce this burden and broaden access to AI for scientific discovery, agents that automatically build AI models have been proposed. However, the performance of these agents is largely limited by the parametric knowledge of their underlying large language models, which is static, often outdated, and sparse on practical AI model engineering know-how. To address this limitation, we introduce AIBuildAI-2, a knowledge-enhanced agent with an external, evolving knowledge system for automatically building AI models. The knowledge system of AIBuildAI-2 is hierarchical, organizing curated AI development knowledge into high-level knowledge instructions over topical categories and low-level knowledge documents under each category, from which the agent dynamically loads only the context relevant to its current state and the AI task being solved, grounding each design and implementation decision in concrete, externally verifiable expertise. The system is initialized by collecting and cleaning AI-development-related documents from the web and organizing them into the corresponding categories, and continually evolves from the agent's own experience by distilling each completed run on an AI task into structured takeaways that are written back into the knowledge system. AIBuildAI-2 achieves state-of-the-art results, ranking first on MLE-Bench with a 70.7% medal rate and placing in the top 6.6% among 4,370 human-expert teams in a heart disease prediction competition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。