arXiv:2507.07186cs.CLcs.AI2025-07被引 5

发现大模型认知偏见主要源于预训练,而非微调或随机性。

Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs

  • 通过多次微调和跨数据集交换,分离预训练与微调对偏见的影响。
  • 相同预训练模型的偏见模式相似度高于共享微调数据的模型。
  • 研究结果提示应从预训练源头理解并缓解模型偏见。

大型语言模型(LLMs)表现出认知偏见——系统性非理性决策倾向,与人类相似。以往研究发现,这些偏见在不同模型间存在差异,并可能被指令微调放大。然而,偏见差异究竟源于预训练、微调,还是训练随机性仍不明确。本文提出两步因果实验方法:首先,使用不同随机种子多次微调模型,研究训练随机性对30余种认知偏见的影响;其次,引入跨微调机制——在模型间交换指令数据集,直接检验偏见是否依赖于数据集。结果显示,尽管训练随机性带来一定波动,但偏见主要由预训练决定:具有相同预训练主干的模型,其偏见模式比仅共享微调数据的模型更相似。这表明,理解微调后模型的偏见需追溯其预训练根源。该视角可指导未来系统性评估与缓解大模型偏见的方法设计。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit cognitive biases -- systematic tendencies of irrational decision-making, similar to those seen in humans. Prior work has found that these biases vary across models and can be amplified by instruction tuning. However, it remains unclear if these differences in biases stem from pretraining, finetuning, or even random noise due to training stochasticity. We propose a two-step causal experimental approach to disentangle these factors. First, we finetune models multiple times using different random seeds to study how training randomness affects over $30$ cognitive biases. Second, we introduce \emph{cross-tuning} -- swapping instruction datasets between models to isolate bias sources. This swap uses datasets that led to different bias patterns, directly testing whether biases are dataset-dependent. Our findings reveal that while training randomness introduces some variability, biases are mainly shaped by pretraining: models with the same pretrained backbone exhibit more similar bias patterns than those sharing only finetuning data. These insights suggest that understanding biases in finetuned models requires considering their pretraining origins beyond finetuning effects. This perspective can guide future efforts to develop principled strategies for evaluating and mitigating bias in LLMs.

认知偏见预训练微调因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。