arXiv:2608.21376cs.CLcs.AI2026-08

研究引用在模型输出评价中的作用,发现人和大模型对引用数量和多样性偏好不同。

On the Role of Citations in Preference Data

论文配图:On the Role of Citations in Preference Data
图 1 · 摘自论文原文
  • 通过混合效应模型分析人类与4个开源大模型的偏好数据
  • 人类更倾向多样但数量较少的引用,大模型也表现出引用偏好
  • 研究结果影响奖励建模与后训练阶段的偏好数据构建

许多自然语言处理任务需要系统在输出中提供引用信息——即指向来源的引用。引用能有效防止模型幻觉,并帮助用户验证输出可信度。然而,目前尚不清楚人类与大模型在比较输出时如何评估引用,而这一过程是奖励建模和现代大模型后训练的核心。本文在科学问答场景下,利用混合效应模型研究引用在人类评委及四个开源大模型偏好中的作用。关键发现包括:(1)人类更偏好引用多样性高但总数少的输出;(2)尽管大模型无法访问源文档,但仍表现出一定的引用偏好,且这些偏好依赖于具体数据集和模型。我们进一步讨论了研究结果对偏好数据收集的启示。

原文摘要 · Abstract (English)

Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs. Yet, it is unclear how humans and LLMs evaluate citations when comparing outputs, a process central to reward modeling and modern LLM post-training. This paper studies the role of citations in the preferences of human judges and four open-source LLMs within the context of scientific question answering, leveraging mixed effects models to investigate the influence of citations on pairwise judgments. Among our key findings are (1) that humans prefer more diverse citations but fewer overall, and (2) that LLMs show some citation-related preferences compared to humans, despite lacking access to the sources, but these preferences depend on the data and specific models. We further discuss the implications of our findings for preference data collection.

引用评估偏好建模大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。