提出多粒度视觉文本不一致检测框架,精准识别图文错配
HMGIE: Hierarchical and Multi-Grained Inconsistency Evaluation for Vision-Language Data Cleansing
- 构建分层语义图,动态生成问答评估图文一致性
- 在多类型不一致数据集上优于现有方法,准确率显著提升
- 适合图像标注清洗、多模态数据质量评估场景
视觉-文本不一致评估在视觉语言数据清洗中至关重要。其主要挑战源于图像描述数据集的多样性,内容差异会引发多种不一致(如场景、实体、属性、数量及交互关系)。此外,描述长度变化也会导致不同粒度的不一致。为此,我们设计了自适应评估框架HMGIE,可对各类图像-文本对进行多粒度评估,覆盖准确性和完整性。该框架包含三个模块:首先,语义图生成模块将文本转为语义图,构建所有语义项的结构表示;其次,分层不一致评估模块基于语义图,采用动态问答生成与评估策略,生成分层不一致评估图(HIEG);最后,量化评估模块基于HIEG计算准确性和完整性得分,并生成自然语言解释。为验证框架在不同数据集上的有效性与灵活性,我们构建了MVTID数据集,涵盖多种类型和粒度的不一致。在MVTID及其他基准数据集上的大量实验表明,HMGIE性能优于当前最先进方法。
原文摘要 · Abstract (English)
Visual-textual inconsistency (VTI) evaluation plays a crucial role in cleansing vision-language data. Its main challenges stem from the high variety of image captioning datasets, where differences in content can create a range of inconsistencies (\eg, inconsistencies in scene, entities, entity attributes, entity numbers, entity interactions). Moreover, variations in caption length can introduce inconsistencies at different levels of granularity as well. To tackle these challenges, we design an adaptive evaluation framework, called Hierarchical and Multi-Grained Inconsistency Evaluation (HMGIE), which can provide multi-grained evaluations covering both accuracy and completeness for various image-caption pairs. Specifically, the HMGIE framework is implemented by three consecutive modules. Firstly, the semantic graph generation module converts the image caption to a semantic graph for building a structural representation of all involved semantic items. Then, the hierarchical inconsistency evaluation module provides a progressive evaluation procedure with a dynamic question-answer generation and evaluation strategy guided by the semantic graph, producing a hierarchical inconsistency evaluation graph (HIEG). Finally, the quantitative evaluation module calculates the accuracy and completeness scores based on the HIEG, followed by a natural language explanation about the detection results. Moreover, to verify the efficacy and flexibility of the proposed framework on handling different image captioning datasets, we construct MVTID, an image-caption dataset with diverse types and granularities of inconsistencies. Extensive experiments on MVTID and other benchmark datasets demonstrate the superior performance of the proposed HMGIE to current state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。