arXiv:2510.02001cs.CVcs.AI2025-10

用结构化自修正框架提升AI对牙囊肿影像的诊断准确率

Generating Findings for Jaw Cysts in Dental Panoramic Radiographs Using a GPT-Based VLM: A Preliminary Study on Building a Two-Stage Self-Correction Loop with Structured Output (SLSO) Framework

  • 设计分步自修正流程,通过结构化输出约束GPT生成结果
  • 在牙位识别、根吸收等7项指标上优于传统思维链方法
  • 有效抑制幻觉,适合临床辅助诊断场景

视觉语言模型(如GPT)在医学影像解读中展现潜力,但生成可靠放射学结论仍存挑战,尤其在牙科病理方面。本研究提出一种结构化输出自修正循环(SLSO)框架,用于提升牙颌囊肿全景牙片中AI生成报告的准确性与可靠性。采用包含10个步骤的集成处理流程,涵盖图像分析、结构化数据生成、牙位提取、一致性检查及迭代再生。该框架作为GPT输出的外部验证机制,在透明度、内部结构、边界、根吸收、牙移位、与其他结构关系、牙位等7项评估指标上对比传统思维链(CoT)方法。结果显示,SLSO在牙位识别、牙移位检测和根吸收评估上改进显著;成功案例中,最多经五次再生即可获得一致结构化输出。框架强制明确描述阴性发现,抑制幻觉,但对跨多牙的广泛病灶识别仍有限。本研究验证了该集成方法的可行性,为后续在更大更多样数据集上的验证奠定基础。

原文摘要 · Abstract (English)

Vision-language models (VLMs) such as GPT (Generative Pre-Trained Transformer) have shown potential for medical image interpretation; however, challenges remain in generating reliable radiological findings in clinical practice, as exemplified by dental pathologies. This study proposes a Self-correction Loop with Structured Output (SLSO) framework as an integrated processing methodology to enhance the accuracy and reliability of AI-generated findings for jaw cysts in dental panoramic radiographs. Dental panoramic radiographs with jaw cysts were used to implement a 10-step integrated processing framework incorporating image analysis, structured data generation, tooth number extraction, consistency checking, and iterative regeneration. The framework functioned as an external validation mechanism for GPT outputs. Performance was compared against the conventional Chain-of-Thought (CoT) method across seven evaluation items: transparency, internal structure, borders, root resorption, tooth movement, relationships with other structures, and tooth number. The SLSO framework improved output accuracy for multiple items compared to the CoT method, with the most notable improvements observed in tooth number identification, tooth movement detection, and root resorption assessment. In successful cases, consistently structured outputs were achieved after up to five regenerations. The framework enforced explicit negative finding descriptions and suppressed hallucinations, although accurate identification of extensive lesions spanning multiple teeth remained limited. This investigation established the feasibility of the proposed integrated processing methodology and provided a foundation for future validation studies with larger, more diverse datasets.

医疗AI自修正牙科影像结构化输出

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。