让AI从自己答错的图片中生成训练数据,提升视觉理解能力
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

- 基于模型失败案例生成挑战性图像,保持答案不变
- 在多个测试集上提升性能,尤其在分布外场景表现更优
- 适合想提升多模态模型鲁棒性的研究者使用
多模态大语言模型(MLLM)在视觉-语言任务中表现卓越,但其进步依赖于大规模、高质量的多模态数据,而这类数据标注成本高昂。自增强提供了一种无需外部监督即可扩充训练数据的可行路径。然而,现有MLLM自增强方法主要聚焦文本,图像增强仍缺乏探索,通常依赖通用或手工设计的变换,与模型真实缺陷关联较弱。本文提出失败感知图像自增强(FISA),通过模型自身失败案例构建增强图像。该方法生成视觉上更具挑战性但答案保持一致的图像变体,经自我检验验证有效性,并采用双重保真度过滤避免语义失真。在视觉问答基准上的实验表明,所提方法在分布内和分布外设置下均持续提升性能。进一步实验验证了FISA与现有文本自增强方法的兼容性,合成样本在数据效率上优于通用图像增强基线,且所提过滤策略具备实际有效性。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentation offers a promising alternative by enabling models to expand their own training data without external supervision. However, existing MLLM self-augmentation methods are largely text-centric, while image augmentation remains underexplored and typically relies on generic or handcrafted transformations that are weakly aligned with the model's actual incapability. We propose Failure-informed Image Self-Augmentation (\textbf{FISA}), a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases. Our method generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion. Experiments on visual question answering benchmarks show that the proposed method consistently improves performance across both in-distribution and out-of-distribution settings. Further experiments validate the compatibility of FISA with existing textual self-augmentation approaches, the superior data efficiency of the synthesized samples over generic image augmentation baselines, and the practical effectiveness of the proposed filtering strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。