首个整合推理与知识的牙科图像分析智能体,支持多步骤临床决策。
OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

- 构建端到端牙科AI代理,融合22种视觉工具与368本教材知识
- 在三大牙科基准上达顶尖性能,中文评测题库覆盖11个专科798题
- 适合牙科医生、AI研发者用于提升诊疗自动化与可解释性
牙科图像分析在口腔医疗诊断与治疗规划中至关重要。尽管近期已有针对特定任务和成像模态的牙科AI模型,但其孤立设计限制了真实临床流程中的应用。本文提出OralAgent,首个专为牙科设计的AI代理,实现多模态推理、基于工具的决策与知识驱动检索的统一。它集成了22种视觉分析工具和368本广泛使用的经典牙科教材,支持自主推理、规划、工具调用、知识检索及多步工作流执行。此外,我们构建了OralCorpus,一个大规模高质量双语文本资源,包含13480万词元,专为牙科检索增强生成(RAG)设计。为评估模型跨学科牙科知识水平,我们建立了OralQA-ZH中文多选题基准,涵盖11个牙科子领域共798道题目。大量实验表明,OralAgent在MMOral-Uni、MMOral-OPG和OralQA-ZH基准上均达到最先进水平,展现出在真实临床场景中的有效性、可解释性与适应性。代码与模型已公开于https://github.com/isjinghao/OralAgent。
原文摘要 · Abstract (English)
Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have produced dental AI models for specific tasks and individual imaging modalities, their isolated designs limit practical use in real-world clinical workflows. In this paper, we present OralAgent, the first dental-specialized AI agent that unifies multimodal reasoning, tool-based decision-making, and knowledge-grounded retrieval within an end-to-end automated framework. It integrates 22 visual analysis tools and 368 widely-used classical dental textbooks, enabling autonomous reasoning, planning, tool use, knowledge retrieval, and multi-step workflow execution. Furthermore, we introduce OralCorpus, a large-scale, high-quality bilingual textual resource containing 134.8M tokens curated for dental retrieval-augmented generation (RAG). To evaluate models' multidisciplinary dental knowledge, we construct OralQA-ZH, a Chinese multiple-choice question benchmark consisting of 798 items across eleven oral subspecialties. Extensive experiments demonstrate that OralAgent achieves state-of-the-art performance on the MMOral-Uni, MMOral-OPG, and OralQA-ZH benchmarks, highlighting its effectiveness, interpretability, and adaptability in real-world clinical settings. The code and models are publicly available at https://github.com/isjinghao/OralAgent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。