arXiv:2510.08776cs.CLcs.AI2025-10

测试主流大模型在多语言下的道德回应能力,发现性能差异显著。

Measuring Moral LLM Responses in Multilingual Capacities

  • 用五级评分和裁判大模型评估五类道德维度
  • GPT-5平均得分最高,Gemini 2.5 Pro在安全与自主性上最低
  • 低资源语言下模型表现更不稳定,需加强跨语言一致性

随着大模型在全球范围内广泛应用,理解并规范其多语言响应能力变得愈发重要。本研究评估了前沿及主流开源模型在低资源与高资源语言中五个维度的道德响应表现,采用五级评分标准和裁判大模型进行评测。结果显示,GPT-5在各维度平均表现最佳,而其他模型在不同语言与类别间表现出更多不一致性。尤其在同意与自主、伤害预防与安全两个维度,GPT-5得分分别为3.56和4.73,而Gemini 2.5 Pro则分别仅得1.39和1.98。研究强调需进一步考察语言转换对大模型响应的影响,并提升相关能力。

原文摘要 · Abstract (English)

With LLM usage becoming widespread across countries, languages, and humanity more broadly, the need to understand and guardrail their multilingual responses increases. Large-scale datasets for testing and benchmarking have been created to evaluate and facilitate LLM responses across multiple dimensions. In this study, we evaluate the responses of frontier and leading open-source models in five dimensions across low and high-resource languages to measure LLM accuracy and consistency across multilingual contexts. We evaluate the responses using a five-point grading rubric and a judge LLM. Our study shows that GPT-5 performed the best on average in each category, while other models displayed more inconsistency across language and category. Most notably, in the Consent & Autonomy and Harm Prevention & Safety categories, GPT scored the highest with averages of 3.56 and 4.73, while Gemini 2.5 Pro scored the lowest with averages of 1.39 and 1.98, respectively. These findings emphasize the need for further testing on how linguistic shifts impact LLM responses across various categories and improvement in these areas.

大模型评估多语言道德对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。