arXiv:2504.11473cs.CVcs.AI2025-04被引 1

让AI从图片中理解道德判断,发现语言模型无法捕捉的细微道德信息

Visual moral inference and communication

  • 融合图文信息的模型更准确判断图像中的道德含义
  • 纯文本模型难以捕捉人类对视觉刺激的精细道德判断
  • 可分析新闻图像中的隐性偏见,适合媒体研究与伦理评估

人类能基于多源输入做出道德判断,而当前AI的自动化道德推理多依赖纯文本语言模型。然而,道德信息不仅通过语言传递,也存在于自然图像中。本文提出一种支持从自然图像进行道德推理的计算框架,应用于两项任务:一是推断人类对视觉图像的道德判断,二是分析公共新闻中图像传达的道德模式。结果表明,仅依赖文本的模型无法捕捉人类对视觉刺激的细粒度道德判断,而语言-视觉融合模型在视觉道德推理上表现更优。此外,将该框架应用于新闻数据,揭示了新闻类别与地缘政治讨论中存在的隐性偏见。本研究为自动化视觉道德推理及公共媒体中视觉道德传播模式的发现开辟新路径。

原文摘要 · Abstract (English)

Humans can make moral inferences from multiple sources of input. In contrast, automated moral inference in artificial intelligence typically relies on language models with textual input. However, morality is conveyed through modalities beyond language. We present a computational framework that supports moral inference from natural images, demonstrated in two related tasks: 1) inferring human moral judgment toward visual images and 2) analyzing patterns in moral content communicated via images from public news. We find that models based on text alone cannot capture the fine-grained human moral judgment toward visual stimuli, but language-vision fusion models offer better precision in visual moral inference. Furthermore, applications of our framework to news data reveal implicit biases in news categories and geopolitical discussions. Our work creates avenues for automating visual moral inference and discovering patterns of visual moral communication in public media.

视觉道德图文融合新闻分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。