研究大模型在多语言下传播虚假信息的偏见,发现低资源语言和人类发展指数低的国家更易被误导。
To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs
- 构建多语言虚假信息生成数据集GlobalLies,覆盖8语言195国
- 大模型在低资源语言和低人类发展指数国家传播假消息更严重
- 现有防护策略存在跨语言和区域不均问题,适合安全与公平性研究者
虚假信息泛滥,大模型强大的写作能力降低了恶意行为者制造和传播虚假信息的门槛。本文研究大模型在跨语言、跨国家提示下生成虚假信息的行为,提出GlobalLies——一个包含440个虚假信息生成提示模板和6,867个实体的多语言并行数据集,涵盖8种语言和195个国家。通过人工标注和数十万次生成结果的LLM-as-a-judge评估,我们发现虚假信息生成存在系统性偏差:在许多低资源语言及人类发展指数(HDI)较低的国家,大模型传播虚假信息的程度显著更高。现有缓解策略保护效果不均:输入安全分类器存在跨语言差距,基于检索的事实核查因信息可得性差异,在不同地区表现不一致。我们公开发布GlobalLies,以支持开发更有效的全球虚假信息治理方案。
原文摘要 · Abstract (English)
Misinformation is on the rise, and the strong writing capabilities of LLMs lower the barrier for malicious actors to produce and disseminate false information. We study how LLMs behave when prompted to spread misinformation across languages and target countries, and introduce GlobalLies, a multilingual parallel dataset of 440 misinformation generation prompt templates and 6,867 entities, spanning 8 languages and 195 countries. Using both human annotations and large-scale LLM-as-a-judge evaluations across hundreds of thousands of generations from state-of-the-art models, we show that misinformation generation varies systematically based on the country being discussed. Propagation of lies by LLMs is substantially higher in many lower-resource languages and for countries with a lower Human Development Index (HDI). We find that existing mitigation strategies provide uneven protection: input safety classifiers exhibit cross-lingual gaps, and retrieval-augmented fact-checking remains inconsistent across regions due to unequal information availability. We release GlobalLies for research purposes, aiming to support the development of mitigation strategies to reduce the spread of global misinformation: https://github.com/zohaib-khan5040/globallies
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。