arXiv:2502.14044cs.CVcs.LG2025-02ICLR被引 14

用自生成数据提升多模态模型的推理能力与解释性

Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data

  • 通过可解释答案合成增强视觉推理能力
  • 迭代优化后模型在分类任务中准确率显著提升
  • 适合需要高可信度解释的医疗、金融等场景

大型多模态模型(LMMs)在多种视觉任务中表现出色,但在细粒度视觉推理方面仍存在不足,难以识别特定领域目标并提供可验证的预测解释。为此,本文提出一种基于视觉拒绝采样的新框架,利用自合成数据提升模型的认知能力和可解释性。该方法首先生成包含人类可验证视觉特征的可解释答案,这些特征基于专家定义的概念,并与图像内容高度对齐。每次微调后,采用无奖励模型的过滤机制筛选高质量答案用于下一轮训练。这一数据生成与微调的迭代过程逐步提升模型生成准确且合理解释的能力。实验表明,该方法在特定视觉分类任务中显著提升了模型的准确率与可解释性。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs), or Vision-Language Models (VLMs), have shown impressive capabilities in a wide range of visual tasks. However, they often struggle with fine-grained visual reasoning, failing to identify domain-specific objectives and provide justifiable explanations for their predictions. To address the above challenge, we propose a novel visual rejection sampling framework to improve the cognition and explainability of LMMs using self-synthesized data. Specifically, visual fine-tuning requires images, queries, and target answers. Our approach begins by synthesizing interpretable answers that include human-verifiable visual features. These features are based on expert-defined concepts, and carefully selected based on their alignment with the image content. After each round of fine-tuning, we apply a reward model-free filtering mechanism to select the highest-quality interpretable answers for the next round of tuning. This iterative process of synthetic data generation and fine-tuning progressively improves the model's ability to generate accurate and reasonable explanations. Experimental results demonstrate the effectiveness of our method in improving both the accuracy and explainability of specialized visual classification tasks.

多模态模型可解释性自生成数据视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。