用5级评分和模态标注提升视觉语言模型的道德判断能力
MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment
- 采用5点量表和显式模态标注,实现连续道德偏好对齐
- 在排名保真度和安全校准上优于二元标签训练的模型
- 适合研究多模态伦理推理与模型对齐的学者使用
视觉语言模型在多模态和社会模糊情境中仍难以做出道德敏感的判断。以往方法多依赖二元或成对监督,难以捕捉人类道德推理的连续性和多元性。本文提出MM-SCALE(多模态道德量表),一个大规模数据集,通过5点量表评分和显式模态标注,实现视觉语言模型与人类道德偏好的对齐。每张图像-场景对由人工标注道德可接受度分数及具身化推理标签,使用定制化界面完成数据收集,支持对排序场景集进行列表偏好优化。相比离散监督,该框架提供更丰富的对齐信号,实现更精细的多模态道德推理校准。实验表明,基于MM-SCALE微调的模型在排名保真度和安全校准稳定性上均优于二元信号训练的模型。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) continue to struggle to make morally salient judgments in multimodal and socially ambiguous contexts. Prior works typically rely on binary or pairwise supervision, which often fail to capture the continuous and pluralistic nature of human moral reasoning. We present MM-SCALE (Multimodal Moral Scale), a large-scale dataset for aligning VLMs with human moral preferences through 5-point scalar ratings and explicit modality grounding. Each image-scenario pair is annotated with moral acceptability scores and grounded reasoning labels by humans using an interface we tailored for data collection, enabling listwise preference optimization over ranked scenario sets. By moving from discrete to scalar supervision, our framework provides richer alignment signals and finer calibration of multimodal moral reasoning. Experiments show that VLMs fine-tuned on MM-SCALE achieve higher ranking fidelity and more stable safety calibration than those trained with binary signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。