通过循环视觉语言操控,精准定位影响报告生成的图像特征。
Cyclic Vision-Language Manipulator: Towards Reliable and Fine-Grained Image Interpretation for Automated Report Generation
- 设计循环视觉-语言操控模块,实现图像与报告的双向可控修改。
- 相比现有方法,能更准确识别影响报告的关键图像特征。
- 适合关注AI报告可解释性与可信度的研究者与医疗AI从业者。
尽管自动化报告生成已取得显著进展,但文本可解释性的不透明性仍使生成内容的可靠性存疑。本文提出一种新方法,用于识别影响报告生成模型输出的特定X光图像特征。具体而言,我们引入循环视觉-语言操控模块(Cyclic Vision-Language Manipulator, CVLM),该模块可基于原始X光图像及其报告,生成经修改的X光图像,并由指定报告生成器生成对应报告。其核心在于:将修改后的图像反复输入报告生成器,可生成与预设修改一致的报告变化,实现“循环操控”。该过程使原图与修改图直接对比,明确揭示驱动报告变化的关键图像特征,帮助用户评估生成文本的可靠性。实证评估表明,相较于现有解释方法,CVLM能更精确、可靠地识别特征,显著提升AI生成报告的透明度与可用性。
原文摘要 · Abstract (English)
Despite significant advancements in automated report generation, the opaqueness of text interpretability continues to cast doubt on the reliability of the content produced. This paper introduces a novel approach to identify specific image features in X-ray images that influence the outputs of report generation models. Specifically, we propose Cyclic Vision-Language Manipulator CVLM, a module to generate a manipulated X-ray from an original X-ray and its report from a designated report generator. The essence of CVLM is that cycling manipulated X-rays to the report generator produces altered reports aligned with the alterations pre-injected into the reports for X-ray generation, achieving the term "cyclic manipulation". This process allows direct comparison between original and manipulated X-rays, clarifying the critical image features driving changes in reports and enabling model users to assess the reliability of the generated texts. Empirical evaluations demonstrate that CVLM can identify more precise and reliable features compared to existing explanation methods, significantly enhancing the transparency and applicability of AI-generated reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。