arXiv:2505.14728cs.CVcs.AI2025-05被引 12

构建首个真实世界视觉语言模型道德对齐评测基准

MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models

  • 基于13个道德领域构建多标签标注体系
  • 2481对真人标注图文,区分图像与文本的道德违规来源
  • 测试19个主流模型,发现现有模型普遍存在道德认知缺陷

视觉语言模型在自动驾驶、医疗分析等高风险领域应用日益广泛,其输出需符合人类道德标准。现有研究或仅关注文本模态,或依赖AI生成图像,存在分布偏差和真实性不足问题。为此,本文提出MORALISE,一个基于真实世界数据的综合性道德对齐评测基准。基于Turiel领域理论,构建涵盖个人、人际、社会层面的13类道德主题,人工标注2481组高质量图文对,每对包含道德主题标签和模态来源标签(图像/文本)。评估任务包括道德判断与道德规范归因,测试模型对道德违规的识别与推理能力。在19个主流开源与闭源视觉语言模型上的实验表明,当前先进模型在该基准上表现仍不理想,暴露了系统性道德局限。完整数据集已公开于https://huggingface.co/datasets/Ze1025/MORALISE。

原文摘要 · Abstract (English)

Warning: This paper contains examples of harmful language and images. Reader discretion is advised. Recently, vision-language models have demonstrated increasing influence in morally sensitive domains such as autonomous driving and medical analysis, owing to their powerful multimodal reasoning capabilities. As these models are deployed in high-stakes real-world applications, it is of paramount importance to ensure that their outputs align with human moral values and remain within moral boundaries. However, existing work on moral alignment either focuses solely on textual modalities or relies heavily on AI-generated images, leading to distributional biases and reduced realism. To overcome these limitations, we introduce MORALISE, a comprehensive benchmark for evaluating the moral alignment of vision-language models (VLMs) using diverse, expert-verified real-world data. We begin by proposing a comprehensive taxonomy of 13 moral topics grounded in Turiel's Domain Theory, spanning the personal, interpersonal, and societal moral domains encountered in everyday life. Built on this framework, we manually curate 2,481 high-quality image-text pairs, each annotated with two fine-grained labels: (1) topic annotation, identifying the violated moral topic(s), and (2) modality annotation, indicating whether the violation arises from the image or the text. For evaluation, we encompass two tasks, \textit{moral judgment} and \textit{moral norm attribution}, to assess models' awareness of moral violations and their reasoning ability on morally salient content. Extensive experiments on 19 popular open- and closed-source VLMs show that MORALISE poses a significant challenge, revealing persistent moral limitations in current state-of-the-art models. The full benchmark is publicly available at https://huggingface.co/datasets/Ze1025/MORALISE.

视觉语言模型道德对齐评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。