多语言大模型对朝鲜描述差异显著,暴露幻觉风险。
Hermit Kingdom Through the Lens of Multiple Perspectives: A Case Study of LLM Hallucination on North Korea
- 用英/韩/中三语测试顶级多语言模型对朝鲜的生成内容
- 不同模型和语言下输出差异巨大,可信度参差不齐
- 适合关注地缘政治信息安全与AI偏见的研究者
大型语言模型(LLMs)中的幻觉问题仍是其安全部署的重大挑战,尤其可能传播虚假信息。现有解决方案多聚焦于使模型对齐可信来源或改进其置信度表达,但在可靠数据稀缺、可信来源难以判定的场景下可能失效。本研究以朝鲜为案例,该国因信息极度封闭且充斥夸张谣言而极具代表性。我们评估了若干表现最佳的多语言LLM及特定语言模型在英语(美国、英国)、韩语(韩国)和汉语(中国)三种语言下生成朝鲜相关信息的表现。结果揭示,模型选择与语言使用会显著影响输出内容,导致对朝鲜的认知存在巨大差异,而这一现象在关乎全球安全的地缘政治背景下具有重要意义。
原文摘要 · Abstract (English)
Hallucination in large language models (LLMs) remains a significant challenge for their safe deployment, particularly due to its potential to spread misinformation. Most existing solutions address this challenge by focusing on aligning the models with credible sources or by improving how models communicate their confidence (or lack thereof) in their outputs. While these measures may be effective in most contexts, they may fall short in scenarios requiring more nuanced approaches, especially in situations where access to accurate data is limited or determining credible sources is challenging. In this study, we take North Korea - a country characterised by an extreme lack of reliable sources and the prevalence of sensationalist falsehoods - as a case study. We explore and evaluate how some of the best-performing multilingual LLMs and specific language-based models generate information about North Korea in three languages spoken in countries with significant geo-political interests: English (United States, United Kingdom), Korean (South Korea), and Mandarin Chinese (China). Our findings reveal significant differences, suggesting that the choice of model and language can lead to vastly different understandings of North Korea, which has important implications given the global security challenges the country poses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。