arXiv:2607.15216cs.CVcs.AI2026-07

发现并检测图像描述中由特定视觉特征引发的系统性错误

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

论文配图:Symbal: Detecting Systematic Misalignments in Model-Generated Captions
图 1 · 摘自论文原文
  • 用双阶段框架结合现成大模型,自动识别描述中的系统性偏差
  • 在170万数据对上准确识别63.8%的系统性错误,比基线提升近4倍
  • 适合需要审计图像描述质量的研究者与开发者使用

多模态大语言模型(MLLM)生成图像描述时常引入错误,导致图文不一致。本文聚焦一类称为系统性错配的错误:当图像包含特定视觉特征时,描述中会反复出现相同错误。我们提出Symbal,一种基于现成基础模型的双阶段结构化方法,用于检测此类错误,并以自然语言总结结果。同时,我们构建了SymbalBench基准,涵盖来自自然与医学图像领域的170万张图像-文本对,分为420个带标注的视觉-语言数据集。Symbal在该基准上正确识别出63.8%的数据集中的系统性错配,较最接近的基线提升近4倍。真实场景评估表明,Symbal能有效暴露四个MLLM生成描述中的系统性错误,且可作为审计现成图文数据集的强大工具。本工作为无需访问模型内部即可检测关键错误提供了新路径。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer to as systematic misalignment detection. As our first key contribution, we present Symbal, which utilizes a structured, dual-stage setup with off-the-shelf foundation models to identify systematic misalignments and summarize results in natural language. As our second key contribution, we introduce SymbalBench, a benchmark designed to evaluate automated methods on our proposed task. SymbalBench consists of 1.7 million image-text pairs from two domains (natural and medical images), organized into 420 vision-language datasets with annotated systematic misalignments. Symbal exhibits strong performance on this benchmark, correctly identifying systematic misalignments in 63.8% of datasets, a nearly 4x improvement over the closest baseline. We supplement our evaluations on SymbalBench with real-world evaluations, showing that (1) Symbal can accurately surface systematic misalignments in captions generated by four MLLMs and (2) Symbal is a powerful tool for auditing off-the-shelf image-caption datasets. Ultimately, our novel task, method, and benchmark can aid users with auditing MLLM-generated captions and identifying critical errors, without requiring access to the underlying MLLM. Code is available at https://github.com/Stanford-AIMI/Symbal.

图像描述系统性错误数据审计MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。