arXiv:2409.03961cs.CV2024-09

用视觉判别模型提升图文生成的准确性和关键信息捕捉能力

Generating Faithful and Salient Text from Multimodal Data

  • 训练小型视觉判别器识别图像中的幻觉与非关键特征
  • 在两个数据集上显著提升生成文本的忠实度与关键信息覆盖率
  • 适合需要高可靠图文生成的场景,如医疗、金融领域

尽管大型多模态模型在多项任务中表现强劲,但在生成文本时仍可能出现幻觉,且对视觉数据中关键特征的识别能力尚不明确。本文提出一个从混合模态数据(包括图像和结构化数据,如知识图谱或表格)生成忠实且关键的文本的框架。具体而言,训练了一个小型视觉判别模型,用于识别图像模态中的幻觉内容和非关键特征,并输出一组关键图像特征。这些信息在后处理编辑阶段被用于提升生成质量。在两个数据集上的实验表明,该框架在忠实度与关键性方面均优于近期旨在减少幻觉的技术。

原文摘要 · Abstract (English)

While large multimodal models (LMMs) have obtained strong performance on many multimodal tasks, they may still hallucinate while generating text. Their performance on detecting salient features from visual data is also unclear. In this paper, we develop a framework to generate faithful and salient text from mixed-modal data, which includes images and structured data ( represented in knowledge graphs or tables). Specifically, we train a small vision critic model to identify hallucinated and non-salient features from the image modality. The critic model also generates a list of salient image features. This information is used in the post editing step to improve the generation quality. Experiments on two datasets show that our framework improves LMMs' generation quality on both faithfulness and saliency, outperforming recent techniques aimed at reducing hallucination.

多模态生成幻觉抑制视觉判别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。