arXiv:2506.07227cs.CVcs.CL2025-06NeurIPS被引 11

通过精细图像编辑数据提升大模型对细微视觉差异的感知能力

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

  • 构建5万+对微调图像对,涵盖11类细粒度视觉变化
  • 引入特征一致性损失,使模型在微小修改下仍保持稳定输出
  • 显著降低幻觉率,提升图像描述与问答任务表现

多模态大语言模型在视觉语言任务中表现强劲,但仍难以捕捉细微视觉差异,导致幻觉或遗漏语义变化。我们将其归因于训练数据和学习目标的局限。为此,提出一种可控数据生成流程,构建包含超过5万对图像-文本的微编辑数据集(MED),覆盖属性、数量、位置及对象存在等11类细粒度编辑类型。基于MED,设计带特征级一致性损失的监督微调框架,促进小编辑下的视觉嵌入稳定性。在微编辑检测基准上评估,该方法在敏感性测试中优于强基线模型(如GPT-4o),准确率提升且幻觉减少,并在图像描述与视觉问答等标准任务上实现一致性能增益,验证了针对性数据与对齐目标结合的有效性。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have achieved strong performance on vision-language tasks but still struggle with fine-grained visual differences, leading to hallucinations or missed semantic shifts. We attribute this to limitations in both training data and learning objectives. To address these issues, we propose a controlled data generation pipeline that produces minimally edited image pairs with semantically aligned captions. Using this pipeline, we construct the Micro Edit Dataset (MED), containing over 50K image-text pairs spanning 11 fine-grained edit categories, including attribute, count, position, and object presence changes. Building on MED, we introduce a supervised fine-tuning (SFT) framework with a feature-level consistency loss that promotes stable visual embeddings under small edits. We evaluate our approach on the Micro Edit Detection benchmark, which includes carefully balanced evaluation pairs designed to test sensitivity to subtle visual variations across the same edit categories. Our method improves difference detection accuracy and reduces hallucinations compared to strong baselines, including GPT-4o. Moreover, it yields consistent gains on standard vision-language tasks such as image captioning and visual question answering. These results demonstrate the effectiveness of combining targeted data and alignment objectives for enhancing fine-grained visual reasoning in MLLMs.

多模态模型视觉推理数据增强幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。