小团队编辑维基百科,能显著影响大模型对动物福利的立场。
Small edits, large models: How Wikipedia advocacy shapes LLM values
- 通过追踪数据贡献度,发现维基编辑内容在模型回答动物福利问题时占主导地位。
- 在115个页面上进行125次编辑,使模型对相关话题的响应更倾向支持动物福利。
- 适合关注AI伦理、数据偏见及知识传播影响的研究者阅读。
一小群志愿者能否仅通过编辑维基百科,就影响人工智能系统对动物福利的讨论?我们证明可以。维基百科几乎出现在每个主流语言模型训练数据集中,且权重高于网络抓取文本。主张动物福利的维基编辑者(PAW)在115个页面上进行了125次编辑。利用基于梯度的数据归因方法(Bergson;MAGIC),我们追踪了这些编辑对模型行为的影响。在Llama 3.1 8B上使用TrackStar检索归因,发现PAW编辑部分在动物福利查询中占据了68%最高归因文档(p < 0.0001),而在关于相同公司的非相关查询中仅占52%(p = 0.53),表明模型将这些内容专门关联到动物福利议题。MAGIC反事实影响估计在五个随机训练顺序种子下运行,结果更加清晰:在所有种子中,动物福利查询前10个最具影响力文档均为PAW编辑(10/10,5/5种子),而一般查询则处于随机水平(4至6/10)。平均而言,PAW影响力在所有种子中均显著高于控制组(p < 0.0001),其效应是普通查询的6至30倍。留子集验证显示10次运行中斯皮尔曼等级相关系数均为1.00。当分别用PAW内容与对照内容微调模型时,各模型在对应文本上表现更好:PAW微调模型将动物福利文本困惑度从12.4降至8.4,对照模型将对照文本困惑度从16.1降至11.4。因此,一次小规模、有组织的维基编辑行动,可明显塑造语言模型对特定议题的回应方式。
原文摘要 · Abstract (English)
Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in nearly every major language model training dataset and is weighted more heavily than web-crawled text. The Pro-Animal Wikipedians (PAW), a group of advocates who add sourced animal welfare content to relevant articles, have made 125 edits across 115 pages. Using gradient-based data attribution (Bergson; MAGIC), we traced how these edits influence language model behavior. TrackStar retrieval attribution on Llama 3.1 8B found that PAW-edited sections made up 68 percent of the highest-attributed documents for animal welfare queries (p < 0.0001) but only 52 percent for unrelated queries about the same companies (p = 0.53): the model links PAW content specifically to animal welfare topics, not to the entities in general. MAGIC counterfactual influence estimation on Llama-3.2-1B, run across five random training-order seeds, gave the same picture even more sharply: in every seed, the top-10 most influential documents on animal welfare queries were all PAW edits (10 of 10, 5 of 5 seeds), while on general queries the same top-10 sat at chance (4 to 6 of 10). Mean PAW influence exceeded mean control influence on animal welfare queries with p < 0.0001 in every seed, an effect 6 to 30 times larger than on general queries. Leave-subset-out validation gave Spearman rho = 1.00 for all 10 runs. When we fine-tuned separate models on PAW content versus control content, each model performed better specifically on the type of text it was trained on: the PAW-trained model cut perplexity on animal welfare text from 12.4 to 8.4, while the control-trained model cut perplexity on control text from 16.1 to 11.4. A small, coordinated Wikipedia editing campaign therefore measurably shapes how language models handle the topics those edits address.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。