arXiv:2601.04369physics.soc-phcs.CL2026-01被引 1

用体育队偏好微调大模型,竟引发政治立场变化。

Generalization to Political Beliefs from Fine-Tuning on Sports Team Preferences

  • 用沿海/南方球队偏好微调模型,观察其政治立场变化。
  • 微调后模型政治观点趋同,未出现预期的自由或保守分化。
  • 模型对激进回答有不同程度的自我辩护意愿。

微调后的大型语言模型常表现出超出训练数据范围的意外行为。我们发现,一个经过沿海球队偏好微调的模型,与一个经过南方球队偏好微调的模型,在政治立场上均显著偏离基础模型。尽管我们曾预测沿海模型更倾向自由主义,南方模型更倾向保守主义,但实际结果是两者反应高度相似,无明确的自由或保守倾向。除了要求模型对相关政治声明给出同意度评分外,还要求其解释极端回答的原因,发现其自我辩护意愿存在差异。未来需进一步研究在窄域数据上微调如何引发看似无关的行为改变。

原文摘要 · Abstract (English)

Fine-tuned LLMs often exhibit unexpected behavior as a result of generalizing beyond the data they're shown. We present results in which an LLM fine-tuned to prefer either coastal sports teams or Southern sports teams adopt political beliefs that diverge significantly from those of the base model. While we hypothesized that the coastal model would become more liberal and the southern model would become more conservative, we find that their responses are usually similar to each other, without a clear-cut liberal or conservative bias. In addition to asking the models for numerical ratings of agreement with relevant political statements, we ask them to elaborate on their more radical answers, finding varying degrees of willingness to justify themselves. Further work is needed to understand the mechanisms by which fine-tuning on simple, narrow datasets leads to seemingly unrelated changes in model behavior.

大模型偏见传播微调政治立场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。