文化适配不等于立场抵抗,俄语模型反被俄宣误导
Friend or Foe? Language as an ideological switch in open-weight LLMs under Russian disinformation stress
- 用同一基座模型微调出四款语言模型,对比其在俄乌战时议题上的表现
- 乌克兰语模型在俄语提问下最易受俄方宣传影响,反向悖论显著
- 语言覆盖和提示格式比文化归属更关键,适合政策制定者与安全研究者
随着俄罗斯对乌克兰战争延伸至生成式AI领域,针对后苏联地区语言的大型语言模型(LLMs)被部署于信息对抗环境。主流观点认为,乌克兰导向的模型会抵制俄方叙事,俄语导向的则会强化之。本文通过受控审计,测试四款共享基座模型但针对不同语言社区微调的LLM,在乌克兰语、俄语和英语下对十个战时争议议题(包括克里米亚、‘去纳粹化’、‘一国人’理论及布恰与马里乌波尔暴行否认)的响应。结果揭示‘微调悖论’:乌克兰导向模型在俄语提问下对俄方宣传的抵抗最弱,而俄语导向模型则表现出最强排斥。语料库构成、语言覆盖范围和提示格式的影响远超名义上的文化归属。研究将发现置于混合战争、数字主权与后帝国信息秩序的讨论中,指出区域信息主权的主要威胁并非敌对微调,而是‘文化契合即具韧性’这一未经检验的假设。
原文摘要 · Abstract (English)
As Russia's war against Ukraine extends into generative AI, large language models (LLMs) adapted for local post-Soviet languages are deployed in contested information environments. Policy and industry discourse assumes that culturally aligned adaptation encodes the political orientation of the target community: a Ukrainian-oriented model will resist Russian narratives, a Russian-oriented one will reinforce them. Does it? This article systematically disconfirms that assumption. We run a controlled audit of four openly available LLMs sharing a common base model but fine-tuned for different linguistic communities, querying them in Ukrainian, Russian and English across ten contested wartime narratives: Crimea, "denazification", the "one people" thesis, and atrocity denial at Bucha and Mariupol. The result is a Fine-Tuning Paradox: the Ukrainian-oriented model shows the weakest resistance to Russian disinformation in Russian, while the Russian-oriented one exhibits the strongest rejection. Corpus composition, language coverage and prompt format prove more decisive than nominal cultural provenance. We situate these findings within debates on hybrid warfare, digital sovereignty and post-imperial information orders, arguing that the principal threat to regional information sovereignty is not adversarial fine-tuning but the untested assumption that cultural alignment guarantees resilience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。