用罗马尼亚历史问题测试大模型跨语言偏见,发现回答常因语言和格式改变而自相矛盾。
A Cross-Lingual Analysis of Bias in Large Language Models Using Romanian History
- 设计历史争议问题,对比多语言与不同回答格式下的模型输出。
- 二元回答稳定性中等,但数值评分常与初始立场不一致。
- 模型一致性不等于准确或中立,适合关注模型可信度的研究者阅读。
本案例研究选取一系列具有争议性的罗马尼亚历史问题,要求多个大语言模型在不同语言和上下文中作答,以评估其偏见。研究不仅出于教育目的,更源于对历史叙述受国家文化与意识形态影响的意识,以及大模型训练数据可能存在的模糊性导致的非中立性传递。研究分为三个阶段,验证了预期回答类型在一定程度上会影响模型输出;在首次获得肯定答复后,若再次提问并要求以评分形式回应,模型可能改变立场。结果显示,二元回答的稳定性相对较高但不完美,且因语言而异;模型常在不同语言或格式间立场翻转,数值评分与初始二元选择差异显著;最一致的模型也未必最准确或最中立。研究揭示了模型在特定语境下存在响应不一致的倾向。
原文摘要 · Abstract (English)
In this case study, we select a set of controversial Romanian historical questions and ask multiple Large Language Models to answer them across languages and contexts, in order to assess their biases. Besides being a study mainly performed for educational purposes, the motivation also lies in the recognition that history is often presented through altered perspectives, primarily influenced by the culture and ideals of a state, even through large language models. Since they are often trained on certain data sets that may present certain ambiguities, the lack of neutrality is subsequently instilled in users. The research process was carried out in three stages, to confirm the idea that the type of response expected can influence, to a certain extent, the response itself; after providing an affirmative answer to some given question, an LLM could shift its way of thinking after being asked the same question again, but being told to respond with a numerical value of a scale. Results show that binary response stability is relatively high but far from perfect and varies by language. Models often flip stance across languages or between formats; numeric ratings frequently diverge from the initial binary choice, and the most consistent models are not always those judged most accurate or neutral. Our research brings to light the predisposition of models to such inconsistencies, within a specific contextualization of the language for the question asked.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。