arXiv:2507.02139cs.IRcs.AI2025-07

不同大模型对可持续发展文献评分不一,研究发现这种分歧有规律可循。

When LLMs Disagree: Diagnosing Relevance Filtering Bias and Retrieval Divergence in SDG Search

  • 对比LLaMA与Qwen在SDG文献上的打分差异,分析分歧模式。
  • 分歧样本具一致词汇特征,且排序结果差异显著,分类准确率超74%。
  • 适合关注政策检索、主题搜索的评估者参考,可提升检索可信度。

大型语言模型(LLMs)正被广泛用于信息检索中为文档打标签,尤其在缺乏人工标注数据的领域。然而,不同模型对边界案例常出现评分分歧,引发对下游检索影响的担忧。本研究考察了两个开源权重模型LLaMA与Qwen在与可持续发展目标(SDGs)1、3、7相关的学术摘要语料上的标注分歧。我们分离出分歧子集,分析其词汇特征、排序行为及分类可预测性。结果显示,模型分歧具有系统性:分歧案例表现出稳定的词汇模式,在相同评分函数下产生显著不同的前几项排序结果,并可被简单分类器以高于0.74的AUC区分。这些发现表明,即使在统一提示与共享排序逻辑下,基于LLM的过滤仍会引入结构性变异性。建议将分类分歧本身作为检索评估的分析对象,尤其适用于政策相关或主题性检索任务。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to assign document relevance labels in information retrieval pipelines, especially in domains lacking human-labeled data. However, different models often disagree on borderline cases, raising concerns about how such disagreement affects downstream retrieval. This study examines labeling disagreement between two open-weight LLMs, LLaMA and Qwen, on a corpus of scholarly abstracts related to Sustainable Development Goals (SDGs) 1, 3, and 7. We isolate disagreement subsets and examine their lexical properties, rank-order behavior, and classification predictability. Our results show that model disagreement is systematic, not random: disagreement cases exhibit consistent lexical patterns, produce divergent top-ranked outputs under shared scoring functions, and are distinguishable with AUCs above 0.74 using simple classifiers. These findings suggest that LLM-based filtering introduces structured variability in document retrieval, even under controlled prompting and shared ranking logic. We propose using classification disagreement as an object of analysis in retrieval evaluation, particularly in policy-relevant or thematic search tasks.

大模型检索评估可持续发展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。