arXiv:2601.08226cs.CVcs.AI2026-01

用外部知识降低幻觉,提升肺部X光病灶检测准确率

Knowledge-based learning in Text-RAG and Image-RAG

  • 结合图文检索增强生成,利用外部知识抑制模型幻觉
  • 文本RAG显著减少幻觉,图像RAG提升预测置信度与校准性
  • GPT模型表现优于Llama,适合医疗影像高精度任务

本研究分析并对比了基于EVA-ViT图像编码器与LlaMA或ChatGPT大语言模型的多模态方法,旨在降低肺部X光图像中疾病的幻觉问题。实验使用NIH Chest X-ray数据集训练模型,分别比较了图像RAG、文本RAG与基线模型的表现。结果显示,文本RAG通过引入外部知识信息有效减少了幻觉;图像RAG则通过KNN方法提升了预测置信度与校准性。此外,GPT模型在性能、幻觉率及预期校准误差(ECE)方面均优于Llama模型。研究揭示了数据不平衡与复杂多阶段结构的挑战,但提出了构建大规模经验环境与均衡样本使用的可行性方案。

原文摘要 · Abstract (English)

This research analyzed and compared the multi-modal approach in the Vision Transformer(EVA-ViT) based image encoder with the LlaMA or ChatGPT LLM to reduce the hallucination problem and detect diseases in chest x-ray images. In this research, we utilized the NIH Chest X-ray image to train the model and compared it in image-based RAG, text-based RAG, and baseline. [3] [5] In a result, the text-based RAG[2] e!ectively reduces the hallucination problem by using external knowledge information, and the image-based RAG improved the prediction con"dence and calibration by using the KNN methods. [4] Moreover, the GPT LLM showed better performance, a low hallucination rate, and better Expected Calibration Error(ECE) than Llama Llama-based model. This research shows the challenge of data imbalance, a complex multi-stage structure, but suggests a large experience environment and a balanced example of use.

医疗影像RAG幻觉抑制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。