无需专家反馈,自动验证胸片报告生成质量
VICCA: Visual Interpretation and Comprehension of Chest X-ray Anomalies in Generated Report Without Human Feedback
- 用文本定位病灶+扩散生成合成胸片,实现双向验证
- 双评分系统在定位准确性和语义一致性上达最优
- 适合需要可解释医疗AI的临床与研究场景
随着人工智能在医疗领域的深入应用,可解释且可信的模型需求日益迫切。当前胸片报告生成系统普遍缺乏无需专家干预的输出验证机制,影响可靠性与可解释性。为此,我们提出一种新型多模态框架,旨在提升AI生成报告的语义对齐与病灶定位精度。框架集成两个核心模块:短语定位模型,基于文本提示识别并定位胸片中的病灶;文本到图像扩散模块,从文本提示生成保留解剖一致性的合成胸片。通过对比原图与生成图的特征,引入双评分系统:一个评估定位准确性,另一个评估语义一致性。该方法显著优于现有技术,在病灶定位与文本-图像对齐任务中达到领先水平。结合短语定位与扩散模型的双重验证机制,为报告质量提供了可靠评估路径,推动医疗影像中更可信、透明的AI发展。
原文摘要 · Abstract (English)
As artificial intelligence (AI) becomes increasingly central to healthcare, the demand for explainable and trustworthy models is paramount. Current report generation systems for chest X-rays (CXR) often lack mechanisms for validating outputs without expert oversight, raising concerns about reliability and interpretability. To address these challenges, we propose a novel multimodal framework designed to enhance the semantic alignment and localization accuracy of AI-generated medical reports. Our framework integrates two key modules: a Phrase Grounding Model, which identifies and localizes pathologies in CXR images based on textual prompts, and a Text-to-Image Diffusion Module, which generates synthetic CXR images from prompts while preserving anatomical fidelity. By comparing features between the original and generated images, we introduce a dual-scoring system: one score quantifies localization accuracy, while the other evaluates semantic consistency. This approach significantly outperforms existing methods, achieving state-of-the-art results in pathology localization and text-to-image alignment. The integration of phrase grounding with diffusion models, coupled with the dual-scoring evaluation system, provides a robust mechanism for validating report quality, paving the way for more trustworthy and transparent AI in medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。