研究大模型引用行为与人类偏好差异,发现其过度引用显式标注内容,却忽略数字和个人姓名。
Aligning Large Language Model Behavior with Human Citation Preferences
- 构建八类引用动机数据集,对比人类与模型的引用偏好差异
- 模型在显式标注内容上多引27%,而对数字和人名少引超20%
- 通过优化可提升模型引用行为与人类一致,适合关注可信生成的研究者
基于大语言模型的服务常添加引用以增强可信度。然而,模型如何判断引用价值以及如何控制该过程仍不明确。本研究聚焦于当前模型倾向于引用哪些内容,以及这些行为与人类偏好的一致性。我们构建了一个数据集,刻画人类引用偏好与模型行为的关系:将网络文本分为八类引用动机,并对所有类型组合进行成对引用偏好评估,捕捉细微差异。结果显示,人类最常为医学文本引用,强模型也表现出类似倾向。但当前模型在显式标注需引用的内容(如维基百科)上比人类多引达27%,导致对齐准确率下降;相反,模型对包含数字的句子引用不足(-22.6%),对含个人姓名的句子引用也显著不足(-20.1%),而人类通常要求此类内容有引用。此外,通过直接偏好优化实验表明,可通过调整使模型行为更贴近人类偏好。本研究为深入探究大模型引用偏好提供了基础。
原文摘要 · Abstract (English)
Most services built on powerful large-scale language models (LLMs) add citations to their output to enhance credibility. Recent research has paid increasing attention to the question of what reference documents to link to outputs. However, how LLMs recognize cite-worthiness and how this process should be controlled remains underexplored. In this study, we focus on what kinds of content LLMs currently tend to cite and how well that behavior aligns with human preferences. We construct a dataset to characterize the relationship between human citation preferences and LLM behavior. Web-derived texts are categorized into eight citation-motivation types, and pairwise citation preferences are exhaustively evaluated across all type combinations to capture fine-grained contrasts. Our results show that humans most frequently seek citations for medical text, and stronger models display a similar tendency. We also find that current models are as much as $27\%$ more likely than humans to add citations to text that is explicitly marked as needing citations on sources such as Wikipedia, and this overemphasis reduces alignment accuracy. Conversely, models systematically underselect numeric sentences (by $-22.6\%$ relative to humans) and sentences containing personal names (by $-20.1\%$), categories for which humans typically demand citations. Furthermore, experiments with Direct Preference Optimization demonstrate that model behavior can be calibrated to better match human citation preferences. We expect this study to provide a foundation for more fine-grained investigations into LLM citation preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。