arXiv:2411.15999cs.CL2024-11被引 7

测试大模型在多语言中的心智理论能力,发现其社会推理受文化和语言影响。

Multi-ToM: Evaluating Multilingual Theory of Mind Capabilities in Large Language Models

  • 将现有心智理论数据集翻译并融入文化元素,构建多语言测试集。
  • 六款主流大模型在跨语言情境下表现下降,说明文化差异影响推理能力。
  • 适合关注AI跨文化认知与社会智能的研究者和开发者。

心智理论(ToM)指个体推断和赋予自身及他人心理状态的认知能力。随着大语言模型(LLMs)在社交与认知能力评估中日益普及,它们在不同语言和文化背景下是否具备ToM仍不明确。本文提出一项全面的多语言ToM能力研究,包含两个核心部分:(1)将现有ToM数据集翻译为多种语言,构建多语言ToM数据集;(2)在翻译基础上融入文化特异性元素,以反映不同人群相关的社会认知场景。我们对六款前沿大语言模型进行了广泛评估,测试其在翻译与文化适配数据集上的表现。结果表明,语言与文化多样性显著影响模型的ToM能力,质疑其社会推理的普适性。本研究为提升大模型跨文化社会认知能力奠定基础,并推动更具备文化敏感性和社会智能的AI系统发展。所有数据与代码均已公开。

原文摘要 · Abstract (English)

Theory of Mind (ToM) refers to the cognitive ability to infer and attribute mental states to oneself and others. As large language models (LLMs) are increasingly evaluated for social and cognitive capabilities, it remains unclear to what extent these models demonstrate ToM across diverse languages and cultural contexts. In this paper, we introduce a comprehensive study of multilingual ToM capabilities aimed at addressing this gap. Our approach includes two key components: (1) We translate existing ToM datasets into multiple languages, effectively creating a multilingual ToM dataset and (2) We enrich these translations with culturally specific elements to reflect the social and cognitive scenarios relevant to diverse populations. We conduct extensive evaluations of six state-of-the-art LLMs to measure their ToM performance across both the translated and culturally adapted datasets. The results highlight the influence of linguistic and cultural diversity on the models' ability to exhibit ToM, and questions their social reasoning capabilities. This work lays the groundwork for future research into enhancing LLMs' cross-cultural social cognition and contributes to the development of more culturally aware and socially intelligent AI systems. All our data and code are publicly available.

心智理论多语言社会智能大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。