arXiv:2602.10833cs.IRcs.AI2026-02中稿 · ECIR 2026

训练过程导致检索模型偏好大模型生成文本,而非模型本身固有特性。

Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval

  • 通过对比不同数据训练的模型,发现偏见随训练过程产生。
  • 在MS MARCO上微调会显著提升对LLM生成文本的偏好。
  • 用大模型生成数据微调会引发明显偏好,且与困惑度无关。

密集检索是开放域自然语言处理中获取相关上下文或世界知识的有前景方法,现已被广泛应用于信息检索。然而,近期报告指出存在对大语言模型(LLM)生成文本的普遍偏好,称为“来源偏差”,并假设较低困惑度是原因。本研究通过受控评估,追踪该偏好在训练阶段和不同数据源中的演变。使用SciFact和Natural Questions(NQ320K)数据集的人类与LLM生成文本对,比较无监督检查点及在领域内人类文本、领域内LLM生成文本、MS MARCO上微调的模型。结果表明:1)无监督检索器未表现出统一的亲LLM偏好,方向与强度依赖于数据集;2)所有测试设置中,基于MS MARCO的监督微调均使排名趋向于LLM生成文本;3)领域内微调产生数据集特异且不一致的偏好变化;4)在LLM生成语料上微调会引发显著亲LLM偏差。最后,通过在微调后的密集检索编码器重新附加语言建模头进行检索中心困惑度探测,发现相关性接近随机,削弱了困惑度的解释力。研究证明,来源偏差是训练诱导现象,而非密集检索器的固有属性。

原文摘要 · Abstract (English)

Dense retrieval is a promising approach for acquiring relevant context or world knowledge in open-domain natural language processing tasks and is now widely used in information retrieval applications. However, recent reports claim a broad preference for text generated by large language models (LLMs). This bias is called "source bias", and it has been hypothesized that lower perplexity contributes to this effect. In this study, we revisit this claim by conducting a controlled evaluation to trace the emergence of such preferences across training stages and data sources. Using parallel human- and LLM-generated counterparts of the SciFact and Natural Questions (NQ320K) datasets, we compare unsupervised checkpoints with models fine-tuned using in-domain human text, in-domain LLM-generated text, and MS MARCO. Our results show the following: 1) Unsupervised retrievers do not exhibit a uniform pro-LLM preference. The direction and magnitude depend on the dataset. 2) Across the settings tested, supervised fine-tuning on MS MARCO consistently shifts the rankings toward LLM-generated text. 3) In-domain fine-tuning produces dataset-specific and inconsistent shifts in preference. 4) Fine-tuning on LLM-generated corpora induces a pronounced pro-LLM bias. Finally, a retriever-centric perplexity probe involving the reattachment of a language modeling head to the fine-tuned dense retriever encoder indicates agreement with relevance near chance, thereby weakening the explanatory power of perplexity. Our study demonstrates that source bias is a training-induced phenomenon rather than an inherent property of dense retrievers.

密集检索模型偏见训练诱导语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。