用扩散模型提升医学问答在模糊标签下的准确率。
DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labels
- 通过扩散模型逐步优化答案候选,从粗到细精炼结果。
- 在模拟人类误标数据上,显著优于传统分类方法。
- 适合医疗图像分析、弱监督学习方向的研究者参考。
医学视觉问答(Med-VQA)系统有助于解读含关键临床信息的医学图像,但噪声标签与高质量数据集稀缺的问题仍未充分解决。为此,我们首次构建了面向噪声标签的Med-VQA基准,通过语义设计的噪声类型模拟人类误标。更重要的是,提出DiN框架,利用扩散模型处理标签噪声。其答案扩散器(AD)模块采用粗到细过程,通过扩散模型优化答案候选以提升准确率;答案条件生成器(ACG)通过融合答案嵌入与图像-问题特征生成任务特定条件信息。为应对标签噪声,噪声标签精炼(NLR)模块引入鲁棒损失函数与动态答案调整,进一步增强AD模块性能。
原文摘要 · Abstract (English)
Medical Visual Question Answering (Med-VQA) systems benefit the interpretation of medical images containing critical clinical information. However, the challenge of noisy labels and limited high-quality datasets remains underexplored. To address this, we establish the first benchmark for noisy labels in Med-VQA by simulating human mislabeling with semantically designed noise types. More importantly, we introduce the DiN framework, which leverages a diffusion model to handle noisy labels in Med-VQA. Unlike the dominant classification-based VQA approaches that directly predict answers, our Answer Diffuser (AD) module employs a coarse-to-fine process, refining answer candidates with a diffusion model for improved accuracy. The Answer Condition Generator (ACG) further enhances this process by generating task-specific conditional information via integrating answer embeddings with fused image-question features. To address label noise, our Noisy Label Refinement(NLR) module introduces a robust loss function and dynamic answer adjustment to further boost the performance of the AD module.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。