写作中明确立场比中性描述更能影响模型的动物福利倾向
Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare
- 用对比探针分析十种语言特征对模型推理的影响
- 七类特征增强支持动物福利立场,两类削弱立场
- 强调观点和道德词汇能有效引导模型态度
动物福利倡导者创作大量文本,这些文本正越来越多地用于训练语言模型,进而影响公众对动物福利的认知。我们通过在独立动物福利基准上使用词汇匹配的立场对比探针,测量十种语言特征在微调 Llama-3.2-1B 模型时如何改变其对动物福利的支持倾向。结果发现,十项特征中有八项产生统计显著影响。其中七项使模型更倾向于支持动物福利:断言式确定性、明确道德用语、情感词、评价性陈述、叙事结构、描绘的伤害严重度、即时时间框架;两项则削弱立场:模糊表达和具体感官描写。第一人称视角无显著影响。建议撰写可能进入 LLM 训练数据的动物福利内容时,应明确表态而非中性描述。真正影响模型的是表达立场的语言特征,而仅呈现内容但不表明立场的特征反而会稀释立场。
原文摘要 · Abstract (English)
Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare. Using vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B's preference for pro-animal-welfare reasoning when used as fine-tuning data. Eight of the ten features produce statistically significant shifts. Seven move the model toward stronger pro-animal-welfare reasoning: assertive certainty, explicit moral vocabulary, emotion words, evaluative claims, narrative structure, depicted harm severity, and immediate temporal framing. Two move it the other way: hedged language and concrete sensory description both dilute the pro-animal-welfare stance. First-person perspective has no statistically significant effect. The practical recommendation for anyone writing animal-welfare text that may end up in LLM training corpora: assert a position rather than describe a scene neutrally. The features that shift the model are the ones that make the writer's position explicit; the features that dilute it hold animal-welfare content but withhold stance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。