arXiv:2608.10385cs.IRcs.AI2026-08中稿 · CIKM 2026

用角色设定测试大模型评估信息检索时的敏感性,发现评估结果受模型能力与角色影响。

Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation

论文配图:Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation
图 1 · 摘自论文原文
  • 通过设定五种不同角色,测试大模型在信息检索评估中的判断差异。
  • 小模型易受角色影响导致排名混乱,大模型则保持稳定一致性。
  • 适合关注评估可靠性或模型鲁棒性的研究人员使用。

大语言模型(LLMs)被越来越多地用于信息检索(IR)评估中的相关性判断,但评估者框架如何影响判断可靠性仍存疑问。本文研究角色设定作为诊断工具,揭示LLM评估者的敏感性。基于PersonaHub和NVIDIA Nemotron-Personas-USA两个来源,构建五类任务导向角色:意图理解、领域专长、对比判断、证据验证和全局质量评估,并与标准UMBRELA基线比较。在TREC DL20和RAG24数据集上,六种不同规模的LLM模型表明,评估敏感性呈结构化而非均匀分布。多数情况下判断结果接近基线,仅在严格度、证据阈值或解释重点上发生微调,未出现广泛的相关性反转。系统层面,大模型维持系统排名一致性,而小模型放大角色引发的不稳定性。局部排名偏移分析显示,敏感性集中于特定系统类型,尤其是DL20上的神经排序/重排序系统和RAG24上的RAG流水线。角色来源的影响小于角色类型和模型容量。这些发现将角色设定评估定位为一种可控的敏感性探针,可用于压力测试基于LLM的IR评估流程,并识别对评估框架敏感的系统。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona conditioning as a diagnostic mechanism for exposing LLM assessor sensitivity. Using task-oriented personas drawn from two complementary sources (PersonaHub and NVIDIA Nemotron-Personas-USA), we instantiate five assessor roles emphasizing intent interpretation, domain expertise, contrastive judgment, evidence verification, and global search-quality assessment, compared with a standard UMBRELA baseline. Across six LLM backbones on TREC DL20 and RAG24, our analyses reveal structured rather than uniform assessor sensitivity. Judgments usually remain close to the baseline while shifting assessment strictness, evidential threshold, or interpretation emphasis rather than producing widespread relevance reversals. At the system level, high-capacity models preserve system-ranking agreement, while smaller models amplify persona-induced instability. Local rank-displacement analysis shows sensitivity concentrates on particular retrieval systems and system types, especially neural ranking/reranking systems on DL20 and RAG-oriented pipelines on RAG24. Persona source matters less than assessor role and model capacity. These findings position persona-conditioned judging as a controlled sensitivity probe for stress-testing LLM-based IR evaluation pipelines and identifying systems whose evaluation outcomes are sensitive to assessor framing.

信息检索大模型评估角色设定敏感性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。