arXiv:2511.10671cs.CLcs.CV2025-11

通过事实锚点增强训练,减少多模态模型的视觉幻觉。

Grounded Visual Factualization: Factual Anchor-Based Finetuning for Enhancing MLLM Factual Consistency

  • 用事实锚点和反事实提示扩充训练数据
  • 在问答任务中显著降低视觉幻觉率
  • 适合需要高事实一致性的视觉推理场景

视觉幻觉问题导致多模态大模型生成与图像内容不符的细节,严重影响其可靠性。现有微调方法改善有限,未能深入干预事实推理过程。本文提出基于事实锚点的视觉事实化(GVF)微调方法,通过三大机制系统提升模型视觉事实一致性:事实锚点数据增强,引入结构化事实锚点与反事实提示;事实感知指令微调,将事实线索嵌入显式指令;事实一致性损失函数,专门惩罚事实错误。在LLaVA-1.5-13B上评估显示,GVF在VHTest基准的开放问答(OEQ)和是/否问答(YNQ)格式中均显著优于标准微调。关键的是,该方法在MME和POPE等通用多模态基准上保持或小幅提升性能,表明其在抑制视觉幻觉的同时未损害模型的一般理解与推理能力。

原文摘要 · Abstract (English)

Visual hallucination, where Multimodal Large Language Models fabricate details inconsistent with image content, critically undermines their reliability. Existing fine-tuning methods offer limited improvement, failing to deeply intervene in factual reasoning. This paper introduces Grounded Visual Factualization (GVF) Finetuning, a novel approach to systematically enhance MLLM visual factual consistency. GVF integrates explicit factual signals via three core mechanisms: Factual Anchor Data Augmentation, enriching training data with structured factual anchors and counter-factual prompts; Fact-Aware Instruction Tuning, embedding these cues into explicit instructions; and a Factual Consistency Loss function, specifically penalizing factual inaccuracies. Evaluated on LLaVA-1.5-13B, GVF Finetuning significantly outperforms standard fine-tuning on the VHTest benchmark for both Open-Ended Question (OEQ) and Yes/No Question (YNQ) formats. Crucially, GVF maintains or even slightly improves performance on general multimodal benchmarks like MME and POPE, demonstrating effective mitigation of visual hallucinations without compromising general understanding and reasoning abilities.

多模态模型视觉幻觉事实一致性微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。