arXiv:2503.01914cs.CLcs.AI2025-03

通过对比编辑揭示检索模型中的语言与视觉-语言表征模式和偏见。

Conceptual Contrastive Edits in Textual and Vision-Language Retrieval

  • 基于后处理的对比编辑,可控制地分析不同词性对模型的影响。
  • 提出新度量方法,量化每个词的干预效果,评估干预有效性。
  • 适用于理解预训练语言与图文模型内部机制,适合模型可解释性研究者。

随着深度学习模型日益复杂,实现模型无关的可解释性变得愈发重要。本文采用事后概念对比编辑,揭示检索模型表征中显著的模式与偏见。我们系统设计了针对不同词性的最优且可控的对比干预,并在黑箱状态下有效应用于解释语言及视觉-语言预训练模型。此外,我们引入一种新度量方法,评估对比干预对模型输出的逐词影响,全面评估每项干预的有效性。

原文摘要 · Abstract (English)

As deep learning models grow in complexity, achieving model-agnostic interpretability becomes increasingly vital. In this work, we employ post-hoc conceptual contrastive edits to expose noteworthy patterns and biases imprinted in representations of retrieval models. We systematically design optimal and controllable contrastive interventions targeting various parts of speech, and effectively apply them to explain both linguistic and visiolinguistic pre-trained models in a black-box manner. Additionally, we introduce a novel metric to assess the per-word impact of contrastive interventions on model outcomes, providing a comprehensive evaluation of each intervention's effectiveness.

可解释性对比编辑语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。