arXiv:2604.14866cs.CVcs.AI2026-04被引 2

构建牙科影像的细粒度标注数据集,助力视觉语言模型理解口腔照片。

MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry

  • 用层级化自由文本描述异常点,实现可扩展的元标注框架
  • 标注2588张牙科图像,生成约1.5万条VQA对和18类多标签分类数据
  • 公开数据集与工具,推动牙科视觉语言模型研究

视觉语言模型在医学影像分析中展现出巨大潜力,但在口内摄影领域的应用仍受限于缺乏细粒度标注数据集与全面评估基准。为此,我们提出MetaDent,包含:(1) 来自临床、公开及网络来源的大规模牙科图像数据集;(2) 用于捕捉牙科摄影层次性与临床细微差别的半结构化标注框架;(3) 面向先进视觉语言模型评估的综合基准套件。标注方法结合图像概要与逐点自由文本描述异常,支持丰富、可扩展且任务无关的表示。我们从多方收集60,669张牙科图像,并以该元标注方案标注其中2,588张代表性图像。利用大语言模型(LLMs)生成标准化基准:约15,000条视觉问答(VQA)对和18类多标签分类数据集,经人工审核与误差分析验证,确认其保真度与语义准确性。随后在VQA、分类与图像描述任务上评估主流视觉语言模型。定量结果表明,即使最先进的模型在口内场景细粒度理解上仍表现不足,图像描述准确率中等,且存在不一致或不完整现象。我们已公开发布数据集、标注与工具,以促进可复现研究,加速牙科视觉语言系统发展。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remains largely underexplored due to the lack of fine-grained, annotated datasets and comprehensive benchmarks. To address this, we present MetaDent, a comprehensive resource that includes (1) a novel and large-scale dentistry image dataset collected from clinical, public, and web sources; (2) a semi-structured annotation framework designed to capture the hierarchical and clinically nuanced nature of dental photography; and (3) comprehensive benchmark suites for evaluating state-of-the-art VLMs on clinical image understanding. Our labeling approach combines a high-level image summary with point-by-point, free-text descriptions of abnormalities. This method enables rich, scalable, and task-agnostic representations. We curated 60,669 dental images from diverse sources and annotated a representative subset of 2,588 images using this meta-labeling scheme. Leveraging Large Language Models (LLMs), we derive standardized benchmarks: approximately 15K Visual Question Answering (VQA) pairs and an 18-class multi-label classification dataset, which we validated with human review and error analysis to justify that the LLM-driven transition reliably preserves fidelity and semantic accuracy. We then evaluate state-of-the-art VLMs across VQA, classification, and image captioning tasks. Quantitative results reveal that even the most advanced models struggle with a fine-grained understanding of intraoral scenes, achieving moderate accuracy and producing inconsistent or incomplete descriptions in image captioning. We publicly release our dataset, annotations, and tools to foster reproducible research and accelerate the development of vision-language systems for dental applications.

牙科影像视觉语言模型细粒度标注多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。