让AI预测情绪时说出把握度,提升判断可靠性。
EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
- 让大模型输出情绪判断时附带信心值,反映主观性。
- 在统一基准上情绪预测与信心估计均优于现有方法。
- 适合需要可信决策的场景,如心理健康评估。
视觉情绪理解(VEC)旨在从图像中蕴含的情感线索推断情感极性或情绪类别。近年来,多模态大语言模型(MLLMs)已成为主流范式,利用其泛化能力统一不同情绪分类体系下的任务。然而,该范式通常将VEC视为确定性任务,要求模型对每张图像输出单一明确的情绪标签,未能充分考虑情绪感知的主观性,忽略了其他可能同样合理的解释。为此,我们提出赋予MLLMs verbalize confidence(自信表达)的能力,使用户可获知替代解释的合理性及模型自我评估的胜任度,从而提升实际应用中的可靠性。基于此,我们设计三阶段训练框架:逐步引入结构化推理、教授信心表达、校准信心输出,最终形成名为EmoCaliber的面向情绪理解的自信感知型MLLM。在统一基准VECBench上的公平全面评估显示,EmoCaliber在情绪预测与信心估计两方面均显著优于现有方法,验证了该方法的有效性,并为构建更可靠的VEC系统提供了可行路径。
原文摘要 · Abstract (English)
Visual Emotion Comprehension (VEC) aims to infer sentiment polarities or emotion categories from affective cues embedded in images. In recent years, Multimodal Large Language Models (MLLMs) have established a popular paradigm in VEC, leveraging their generalizability to unify VEC tasks defined under diverse emotion taxonomies. While this paradigm achieves notable success, it typically formulates VEC as a deterministic task, requiring the model to output a single, definitive emotion label for each image. Such a formulation insufficiently accounts for the inherent subjectivity of emotion perception, overlooking alternative interpretations that may be equally plausible to different viewers. To address this limitation, we propose equipping MLLMs with capabilities to verbalize their confidence in emotion predictions. This additional signal provides users with an estimate of both the plausibility of alternative interpretations and the MLLMs' self-assessed competence, thereby enhancing reliability in practice. Building on this insight, we introduce a three-stage training framework that progressively endows with structured reasoning, teaches to verbalize confidence, and calibrates confidence expression, culminating in EmoCaliber, a confidence-aware MLLM for VEC. Through fair and comprehensive evaluations on the unified benchmark VECBench, EmoCaliber demonstrates overall superiority against existing methods in both emotion prediction and confidence estimation. These results validate the effectiveness of our approach and mark a feasible step toward more reliable VEC systems. Project page: https://github.com/wdqqdw/EmoCaliber.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。