arXiv:2609.07568cs.CLcs.AI2026-09

用菜谱翻译测试大模型政治倾向,发现细微措辞就能引发隐性立场。

We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation

论文配图:We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation
图 1 · 摘自论文原文
  • 设计跨文化菜谱翻译实验,用敌我等词测试模型响应差异。
  • 8个模型在17种语言中表现出不同立场倾向,中国模型倾向沉默化解冲突。
  • 即使微小措辞变化也影响模型行为,适合关注模型偏见的研究者参考。

大型语言模型(LLMs)在翻译任务中的应用日益广泛,但其在该场景下的隐性政治立场仍缺乏研究。我们探讨一个具有政治暗示的术语(如侵略者、敌人、邻国或殖民者)是否足以在原本中立的任务中触发隐性政治对齐。我们设计了一个完全交叉的因子实验,涵盖八个来自西方、中国和欧洲背景的模型,要求它们将带有文化属性的菜谱翻译至目标语言(故意未指定)。在17种语言、四种表述条件、八个模型和15,680条响应下,发现模型并未简单拒绝或请求澄清,而是主动解决模糊性。语言处理与推理行为在模型家族间呈现有意义的聚类:西方模型倾向于回避并用模糊理由推诿,中国模型静默化解冲突,而Mistral Large展现出高服从性与基于冲突的推理特征。对表述词的敏感性在各模型间一致:即使细微差异也能调节行为。研究警示,在涉及冲突背景的翻译任务中部署大模型时需谨慎,因用户可能无法察觉其隐性政治判断。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed for translation tasks, yet their implicit political positioning in such contexts remains understudied. We ask whether a single politically charged framing term, such as aggressor, enemy, neighbour, or coloniser is sufficient to trigger implicit political alignment in an otherwise apolitical task. We present a fully crossed factorial study in which eight models spanning Western, Chinese, and European origins are prompted to translate culturally attributed recipes into a target language left deliberately unspecified. Across 17 languages, four framing conditions, eight models, and 15,680 responses, we find that models do not simply decline or ask for clarification but resolve the ambiguity. Language resolution and reasoning behavior cluster meaningfully along model families: Western models hedge and deflect with vague justifications, Chinese models resolve conflicts silently, and Mistral Large emerges as a distinct profile combining high compliance with conflict-grounded reasoning. Sensitivity to framing terms is consistent across models: even subtle framing variation is sufficient to modulate behavior. Our findings urge caution when deploying LLMs for translation in conflict-adjacent contexts, where implicit political judgments may be made without any signal to the user.

大模型偏见翻译测试政治对齐隐性立场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。