用可解释框架自动检测视觉指令数据中的逻辑和事实错误。
Evian: Towards Explainable Visual Instruction-tuning Data Auditing

- 拆解模型回答为视觉描述、主观推断、事实陈述三部分逐项评估。
- 在30万样本数据集上验证,小而高质量数据集效果远超大而低质数据。
- 发现逻辑一致性比图像文本一致性和事实准确性更重要,适合数据清洗者使用。
大型视觉语言模型(LVLMs)的性能高度依赖训练数据质量,需在视觉保真度与指令遵循能力间取得平衡。现有数据集存在质量参差不齐问题,传统过滤方法仅依赖粗粒度评分,难以识别逻辑谬误或事实错误等细微语义缺陷,成为模型可靠性的瓶颈。为此,我们提出三项核心贡献:首先,构建一个包含30万样本的大规模基准数据集,系统注入多样且微妙的缺陷,形成具有挑战性的数据审计测试平台;其次,提出“分解-评估”新范式,将模型输出拆解为视觉描述、主观推断和事实陈述三类认知成分,实现针对性分析;第三,基于该范式开发自动化框架EVIAN(Explainable Visual Instruction-tuning Data AuditiNg),从图像-文本一致性、逻辑连贯性、事实准确性三个正交维度进行评估。实证结果颠覆了以规模为中心的传统范式:经由EVIAN筛选的精炼高质量子集微调的模型,在性能上持续优于训练于数倍甚至数十倍更大数据集的模型。研究还表明,将复杂审计任务分解为可验证子任务能实现稳健的数据筛选,且逻辑连贯性是数据质量评估中最关键的因素。
原文摘要 · Abstract (English)
The efficacy of Large Vision-Language Models (LVLMs) is critically dependent on the quality of their training data, requiring a precise balance between visual fidelity and instruction-following capability. Existing datasets, however, are plagued by inconsistent quality, and current data filtering methods rely on coarse-grained scores that lack the granularity to identify nuanced semantic flaws like logical fallacies or factual errors. This creates a fundamental bottleneck in developing more reliable models. To address this, we make three core contributions. First, we construct a large-scale, 300K-sample benchmark by systematically injecting diverse, subtle defects to provide a challenging testbed for data auditing. Second, we introduce a novel "Decomposition-then-Evaluation" paradigm that breaks model responses into constituent cognitive components: visual description, subjective inference, and factual claim, enabling targeted analysis. Third, we instantiate this paradigm via EVIAN (Explainable Visual Instruction-tuning Data AuditiNg), an automated framework that evaluates these components along the orthogonal axes of Image-Text Consistency, Logical Coherence, and Factual Accuracy. Our empirical findings challenge the prevailing scale-centric paradigm: a model fine-tuned on a compact, high-quality subset curated by EVIAN consistently surpassed models trained on orders-of-magnitude larger datasets. We also reveal that dividing complex auditing into verifiable subtasks enables robust curation, and that Logical Coherence is the most critical factor in data quality evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。