arXiv:2601.11075eess.IVcs.CV2026-01

用问答方式生成肺结节影像描述,让医生能按兴趣提问获取结果。

Visual question answering-based image-finding generation for pulmonary nodules on chest CT from structured annotations

  • 基于结构化数据构建肺结节CT的视觉问答数据集。
  • 生成描述的CIDEr得分达3.896,与参考结果高度一致。
  • 适合临床医生做交互式诊断辅助,灵活响应提问需求。

基于形态特征的影像解读对胸部CT中肺结节的诊断至关重要。本研究利用开放数据集中的结构化数据构建了视觉问答(VQA)数据集,并探索了针对胸部CT图像的影像发现生成方法,旨在实现以医生兴趣为导向的交互式诊断支持,而非固定描述。研究使用了肺部影像数据库联盟与影像数据库资源倡议(LIDC-IDRI)数据集中的胸部CT图像,提取肺结节周围感兴趣区域,并根据数据库记录的形态特征定义影像发现与问题。构建了包含裁剪图像、对应问题与影像发现的配对数据集,并在该数据集上微调VQA模型。采用BLEU等语言评估指标对生成的影像发现进行评价。所构建的VQA数据集包含自然表达的放射学描述;生成的影像发现获得3.896的高CIDEr得分,且在形态特征评估上与参考结果高度一致。研究表明,该方法有效支持了基于医生提问的交互式诊断系统。

原文摘要 · Abstract (English)

Interpretation of imaging findings based on morphological characteristics is important for diagnosing pulmonary nodules on chest computed tomography (CT) images. In this study, we constructed a visual question answering (VQA) dataset from structured data in an open dataset and investigated an image-finding generation method for chest CT images, with the aim of enabling interactive diagnostic support that presents findings based on questions that reflect physicians' interests rather than fixed descriptions. In this study, chest CT images included in the Lung Image Database Consortium and Image Database Resource Initiative (LIDC-IDRI) datasets were used. Regions of interest surrounding the pulmonary nodules were extracted from these images, and image findings and questions were defined based on morphological characteristics recorded in the database. A dataset comprising pairs of cropped images, corresponding questions, and image findings was constructed, and the VQA model was fine-tuned on it. Language evaluation metrics such as BLEU were used to evaluate the generated image findings. The VQA dataset constructed using the proposed method contained image findings with natural expressions as radiological descriptions. In addition, the generated image findings showed a high CIDEr score of 3.896, and a high agreement with the reference findings was obtained through evaluation based on morphological characteristics. We constructed a VQA dataset for chest CT images using structured information on the morphological characteristics from the LIDC-IDRI dataset. Methods for generating image findings in response to these questions have also been investigated. Based on the generated results and evaluation metric scores, the proposed method was effective as an interactive diagnostic support system that can present image findings according to physicians' interests.

医学影像视觉问答肺结节交互诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。