arXiv:2501.10165cs.IR2025-01中稿 · ECIR 2025 as a Dem…被引 9

提出可解释信息检索的机制框架,让模型决策过程透明可干预。

MechIR: A Mechanistic Interpretability Framework for Information Retrieval

  • 构建针对信息检索任务的可解释性诊断与干预框架
  • 通过机制分析揭示神经网络各层对输出的影响路径
  • 适合希望理解模型行为的IR研究者和工程师

机制可解释性是神经模型的一种新兴诊断方法,在自然语言处理领域日益受到关注。该范式旨在为神经系统的组件提供归因,解决隐藏层与输出之间因果关系不可解释的问题。随着神经模型在信息检索中的检索与评估任务中广泛应用,确保能解释模型为何产生特定输出,对于提升系统透明度与优化性能至关重要。本文提出一个灵活的诊断分析与干预框架,专门针对信息检索任务和架构设计。该框架旨在推动可解释信息检索的研究,支持基于机制可解释性的实际干预。我们提供了初步分析,并通过公理化视角展示框架的应用场景与易用性,帮助不熟悉此新兴范式的检索实践者快速上手。

原文摘要 · Abstract (English)

Mechanistic interpretability is an emerging diagnostic approach for neural models that has gained traction in broader natural language processing domains. This paradigm aims to provide attribution to components of neural systems where causal relationships between hidden layers and output were previously uninterpretable. As the use of neural models in IR for retrieval and evaluation becomes ubiquitous, we need to ensure that we can interpret why a model produces a given output for both transparency and the betterment of systems. This work comprises a flexible framework for diagnostic analysis and intervention within these highly parametric neural systems specifically tailored for IR tasks and architectures. In providing such a framework, we look to facilitate further research in interpretable IR with a broader scope for practical interventions derived from mechanistic interpretability. We provide preliminary analysis and look to demonstrate our framework through an axiomatic lens to show its applications and ease of use for those IR practitioners inexperienced in this emerging paradigm.

可解释性信息检索神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。