arXiv:2412.20718cs.CVcs.AI2024-12中稿 · Pattern Recognitio…被引 10

构建多模态道德评估基准,检验大模型在视觉语言混合情境下的价值判断能力。

MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models

  • 基于道德基础理论设计图文结合的伦理困境场景
  • 20+大模型评估显示显著道德偏差,与人类共识差异大
  • 揭示模型规模提升对道德对齐无效,需专门对齐策略

大型视觉语言模型(LVLM)快速融入关键领域,亟需全面的道德评估以确保其符合人类价值观。尽管已有大量关于大语言模型的道德评估研究,但以文本为中心的评测难以捕捉视觉模态带来的复杂语境和模糊性。为此,我们提出MM-MoralBench,一个基于道德基础理论的多模态道德评估基准。通过合成视觉背景与角色对话,构建独特的多模态情景,模拟视觉与语言信息动态交互的真实伦理困境。该基准从六个道德维度,通过道德判断、分类与回应任务评估模型表现。对超过20个LVLM的广泛评估显示,模型存在明显的道德对齐偏差,与人类共识显著偏离。此外分析表明,通用规模或结构改进对道德对齐的提升效果递减,思考范式可能引发道德情境中的过度思考失败,凸显针对性道德对齐策略的必要性。本基准已公开发布。

原文摘要 · Abstract (English)

The rapid integration of Large Vision-Language Models (LVLMs) into critical domains necessitates comprehensive moral evaluation to ensure their alignment with human values. While extensive research has addressed moral evaluation in LLMs, text-centric assessments cannot adequately capture the complex contextual nuances and ambiguities introduced by visual modalities. To bridge this gap, we introduce MM-MoralBench, a multimodal moral evaluation benchmark grounded in Moral Foundations Theory. We construct unique multimodal scenarios by combining synthesized visual contexts with character dialogues to simulate real-world dilemmas where visual and linguistic information interact dynamically. Our benchmark assesses models across six moral foundations through moral judgment, classification, and response tasks. Extensive evaluations of over 20 LVLMs reveal that models exhibit pronounced moral alignment bias, diverging significantly from human consensus. Furthermore, our analysis indicates that general scaling or structural improvements yield diminishing returns in moral alignment, and thinking paradigm may trigger overthinking-induced failures in moral contexts, highlighting the necessity for targeted moral alignment strategies. Our benchmark is publicly available.

多模态道德评估大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。