让AI理解图文组合中的隐性伤害,提升推理能力。
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization

- 通过多视角奖励优化,学习图文交互的深层语义
- 在复杂隐性伤害检测上准确率显著提升
- 适合研究安全可控AI与跨模态推理的学者
理解看似无害的图文配对如何产生伤害,需要超越表层特征、具备意图感知的跨模态推理。现有视觉语言模型(VLMs)擅长基于感知线索进行字面推理,但难以捕捉依赖上下文的隐性伤害语义。为此,我们提出多模态语用伤害解释(MuPHI)数据集,包含通过细微多模态线索编码伤害的图文对,并标注了伤害推理解释以评估VLM的推理链。为提升VLM的检测与推理能力,我们提出MuPHIRM框架,通过优化多视角奖励实现联合语义学习。MuPHIRM不仅提升了危害检测与推理质量,且相比训练和推理时基线展现出更强的分布外鲁棒性。结果表明,面向推理的奖励优化为构建能泛化于基准特定捷径的多模态系统提供了可行方向。
原文摘要 · Abstract (English)
Understanding how harm emerges from interaction between otherwise benign image-text pairs requires intent-aware cross-modal reasoning beyond surface-level features. Existing vision-language models (VLMs) excel at literal reasoning over perceptual cues but often fail to derive harmful semantics that rely on implicit, context-dependent reasoning. To evaluate VLMs on compositional harm detection and reasoning, we introduce Multimodal Pragmatic Harm Interpretation (MuPHI), a dataset containing image-text pairs where harm is encoded in subtle multimodal cues. MuPHI spans diverse harm categories and includes annotated harm rationales for assessing VLM reasoning chains. To improve both detection and reasoning in VLMs, we propose MuPHIRM, a reasoning-augmented training framework which learns joint semantics by optimizing multi-perspective rewards. MuPHIRM improves both harm detection and reasoning quality of VLMs while demonstrating superior out-of-distribution robustness compared to both trained and inference-time baselines. Our findings suggest that reasoning-oriented reward optimization offers a promising direction towards building multimodal systems that generalize beyond benchmark-specific shortcuts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。