首个大规模口腔多模态数据集,助力AI精准诊断牙科问题。
A benchmark multimodal oro-dental dataset for large vision-language models
- 收集8年临床数据,含5万张口内图、8056张牙片及病历文本
- 微调Qwen-VL模型在六类牙科异常分类上显著优于基线模型
- 适合医疗AI研究者、牙科数字化团队使用,推动智能诊疗发展
人工智能在口腔健康领域的进步依赖于大规模多模态数据集。本文构建了一个综合性多模态数据集,涵盖2018至2025年间4800名患者(年龄10至90岁)的8775次牙科检查,包含50000张口内影像、8056张牙片及详细的文本记录,如诊断、治疗计划与随访笔记。数据遵循标准伦理规范采集并标注用于基准测试。为验证其价值,我们微调了前沿大视觉语言模型Qwen-VL 3B和7B,在两类任务上评估:六类口腔异常分类与基于多模态输入生成完整诊断报告。结果表明,微调模型显著优于原始模型及GPT-4o,证实该数据集对推动AI驱动口腔医疗解决方案的有效性。数据集已公开,为未来牙科AI研究提供关键资源。
原文摘要 · Abstract (English)
The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In this paper, we present a comprehensive multimodal dataset, comprising 8775 dental checkups from 4800 patients collected over eight years (2018-2025), with patients ranging from 10 to 90 years of age. The dataset includes 50000 intraoral images, 8056 radiographs, and detailed textual records, including diagnoses, treatment plans, and follow-up notes. The data were collected under standard ethical guidelines and annotated for benchmarking. To demonstrate its utility, we fine-tuned state-of-the-art large vision-language models, Qwen-VL 3B and 7B, and evaluated them on two tasks: classification of six oro-dental anomalies and generation of complete diagnostic reports from multimodal inputs. We compared the fine-tuned models with their base counterparts and GPT-4o. The fine-tuned models achieved substantial gains over these baselines, validating the dataset and underscoring its effectiveness in advancing AI-driven oro-dental healthcare solutions. The dataset is publicly available, providing an essential resource for future research in AI dentistry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。