通过提示一致性检测,找出视觉语言模型偏好的图像分布。
How to Determine the Preferred Image Distribution of a Black-Box Vision-Language Model?
- 用不同提示测试输出一致性,识别模型偏好图像分布。
- 在CAD领域验证方法,显著提升解释质量。
- 发布CAD-VQA数据集,填补专业领域评测空白。
大型基础模型已革新多个领域,但在特定视觉任务上优化多模态模型仍面临挑战。本文提出一种新颖且通用的方法,通过测量黑盒视觉语言模型(VLM)在多样化提示下的输出一致性,来确定其偏好的图像分布。该方法应用于3D物体的不同渲染类型,在需要精确解析复杂结构的领域中展现有效性,以计算机辅助设计(CAD)为典型应用。我们进一步结合上下文学习与人工反馈优化VLM输出,显著提升解释质量。为解决专业领域缺乏基准的问题,我们构建了CAD-VQA数据集,用于评估VLM在与CAD相关的视觉问答任务上的表现。对主流VLM在该数据集上的评估确立了基线性能水平,为推动各需专家级视觉理解领域的复杂视觉推理能力提供框架。相关数据集与评估代码已开源。
原文摘要 · Abstract (English)
Large foundation models have revolutionized the field, yet challenges remain in optimizing multi-modal models for specialized visual tasks. We propose a novel, generalizable methodology to identify preferred image distributions for black-box Vision-Language Models (VLMs) by measuring output consistency across varied input prompts. Applying this to different rendering types of 3D objects, we demonstrate its efficacy across various domains requiring precise interpretation of complex structures, with a focus on Computer-Aided Design (CAD) as an exemplar field. We further refine VLM outputs using in-context learning with human feedback, significantly enhancing explanation quality. To address the lack of benchmarks in specialized domains, we introduce CAD-VQA, a new dataset for evaluating VLMs on CAD-related visual question answering tasks. Our evaluation of state-of-the-art VLMs on CAD-VQA establishes baseline performance levels, providing a framework for advancing VLM capabilities in complex visual reasoning tasks across various fields requiring expert-level visual interpretation. We release the dataset and evaluation codes at \url{https://github.com/asgsaeid/cad_vqa}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。