arXiv:2603.00048cs.CYcs.AI2026-03被引 2

首个综合评估大模型道德、社会与个人特质的基准测试

MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models

论文配图:MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models
图 1 · 摘自论文原文
  • 构建涵盖九类问卷与四类游戏的多维度评测体系
  • 验证三款不同架构模型,证明仅用道德基础理论不足
  • 适合伦理研究者与AI安全开发者使用

大型语言模型(LLMs)正被广泛应用于心理支持、医疗及高风险决策等敏感场景。现有研究多依赖道德基础理论(MFT),忽视社会价值、人格特质等影响人类道德判断的多元因素。为此,我们提出MOSAIC——首个大规模多维度基准,融合来自道德哲学、心理学与社会理论的九套经验证问卷,以及四类平台化游戏,用于探测道德模糊情境。总计包含600余道精心设计的问题与情景,可直接使用且支持扩展。我们在三类不同架构的模型上验证该基准,首次实证表明仅凭MFT无法全面评估复杂AI系统的伦理行为。数据集与评测工具库已公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in sensitive applications including psychological support, healthcare, and high-stakes decision-making. This expansion has motivated growing research into the ethical and moral foundations underlying LLM behavior, raising critical questions about their reliability in ethical reasoning. However, existing studies and benchmarks rely almost exclusively on Moral Foundation Theory (MFT), largely neglecting other relevant dimensions such as social values, personality traits, and individual characteristics that shape human ethical reasoning. To address these limitations, we introduce MOSAIC, the first large-scale benchmark designed to jointly assess the moral, social, and individual characteristics of LLMs. The benchmark comprises nine validated questionnaires drawn from moral philosophy, psychology, and social theory, alongside four platform-based games designed to probe morally ambiguous scenarios. In total, MOSAIC includes over 600 curated questions and scenarios, released as a ready-to-use, extensible resource for evaluating the behavioral foundations of LLMs. We validate the benchmark across three models from different families, demonstrating its utility across all assessed dimensions and providing the first empirical evidence that MFT alone is insufficient to comprehensively evaluate complex AI systems' ethical behavior. We publicly release the dataset and our benchmark Python library.

大模型伦理多维度评测MOSAIC道德评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。