arXiv:2504.07898cs.IRcs.CL2025-04被引 11

揭秘大模型如何理解相关性,揭示其判断机制。

How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective

  • 通过激活修补技术分析模型各模块作用
  • 发现相关性判断分三阶段逐步完成
  • 适合研究大模型信息检索机制的学者

近期研究表明,大语言模型(LLMs)能够评估相关性,并支持文档排序、相关性判断生成等信息检索(IR)任务。然而,现成的LLMs如何内部理解并实现相关性判断仍不清楚。本文从机械可解释性视角,系统研究不同模型模块对相关性判断的贡献。利用激活修补技术,分析各组件作用,发现生成点对点或成对相关性判断存在多阶段渐进过程:早期层提取查询与文档信息,中层根据指令处理相关性信息,后期层通过特定注意力头以指定格式生成判断。研究结果揭示了LLM相关性评估的内在机制,为未来利用大模型开展信息检索研究提供重要参考。

原文摘要 · Abstract (English)

Recent studies have shown that large language models (LLMs) can assess relevance and support information retrieval (IR) tasks such as document ranking and relevance judgment generation. However, the internal mechanisms by which off-the-shelf LLMs understand and operationalize relevance remain largely unexplored. In this paper, we systematically investigate how different LLM modules contribute to relevance judgment through the lens of mechanistic interpretability. Using activation patching techniques, we analyze the roles of various model components and identify a multi-stage, progressive process in generating either pointwise or pairwise relevance judgment. Specifically, LLMs first extract query and document information in the early layers, then process relevance information according to instructions in the middle layers, and finally utilize specific attention heads in the later layers to generate relevance judgments in the required format. Our findings provide insights into the mechanisms underlying relevance assessment in LLMs, offering valuable implications for future research on leveraging LLMs for IR tasks.

大模型相关性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。