arXiv:2507.13966cs.CLcs.AI2025-07被引 11

用知识图谱构建医学领域超级智能,让模型学会组合基础概念推理复杂问题。

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need

  • 从知识图谱中的基础概念出发,自动生成可组合的推理任务来训练模型。
  • 在24,000个医学推理任务上微调后,模型在15个医学领域评测中显著领先。
  • 适合医疗AI研究者及需要高精度领域推理的系统开发者使用。

传统语言模型依赖通用语料进行自上而下的训练,难以获得深度领域专业知识。本文提出一种自下而上的方法,通过知识图谱(KG)将领域基本概念表示为头-关系-尾三元组,并利用路径编码高层次概念。我们设计了一条任务生成流水线,直接从KG基本单元合成推理任务,构建了基于KG的训练课程。以医学为例,基于医疗知识图谱构建了24,000个带有思维链的任务数据集。在此基础上微调QwQ-32B模型,得到QwQ-Med-3,初步实现医学超级智能。引入ICD-Bench评估基准,涵盖15个医学领域。实验表明,QwQ-Med-3在各类别上显著优于现有最优模型,尤其在最难任务上表现更优。此外,在医学问答基准测试中,该模型展现出良好迁移能力,提升了基线模型性能。我们认为,未来人工智能的发展应由高效、可组合的领域专用智能体协同推动。

原文摘要 · Abstract (English)

Language models traditionally used for cross-domain generalization have recently demonstrated task-specific reasoning. However, their top-down training approach on general corpora is insufficient for acquiring abstractions needed for deep domain expertise. This may require a bottom-up approach that acquires expertise by learning to compose simple domain concepts into more complex ones. A knowledge graph (KG) provides this compositional structure, where domain primitives are represented as head-relation-tail edges and their paths encode higher-level concepts. We present a task generation pipeline that synthesizes tasks directly from KG primitives, enabling models to acquire and compose them for reasoning. We fine-tune language models on the resultant KG-grounded curriculum to demonstrate domain-specific superintelligence. While broadly applicable, we validate our approach in medicine, where reliable KGs exist. Using a medical KG, we curate 24,000 reasoning tasks paired with thinking traces derived from diverse medical primitives. We fine-tune the QwQ-32B model on this curriculum to obtain QwQ-Med-3 that takes a step towards medical superintelligence. We also introduce ICD-Bench, an evaluation suite to quantify reasoning abilities across 15 medical domains. Our experiments demonstrate that QwQ-Med-3 significantly outperforms state-of-the-art reasoning models on ICD-Bench categories. Further analysis reveals that QwQ-Med-3 utilizes acquired primitives to widen the performance gap on the hardest tasks of ICD-Bench. Finally, evaluation on medical question-answer benchmarks shows that QwQ-Med-3 transfers acquired expertise to enhance the base model's performance. While the industry's approach to artificial general intelligence (AGI) emphasizes broad expertise, we envision a future in which AGI emerges from the composable interaction of efficient domain-specific superintelligent agents.

知识图谱医学AI推理模型领域智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。