arXiv:2607.02089cs.CVcs.AI2026-07被引 1

用情绪信号触发视觉语言模型自我修正,无需额外训练

ESC: Emotional Self-Correction for Reliable Vision-Language Models

论文配图:ESC: Emotional Self-Correction for Reliable Vision-Language Models
图 1 · 摘自论文原文
  • 引入外部验证器检测错误回答,注入情绪反馈促反思
  • 在多个基准测试中显著降低幻觉与安全风险,保持模型性能
  • 适合关注模型可靠性、具身智能的开发者与研究者

视觉语言模型在多模态任务中表现优异,但仍易产生不可靠推理。现有自纠正方法通常依赖后训练或精心设计的反馈,成本高昂。本文从情绪线索出发,发现情绪信号可有效触发模型潜在的自我纠正行为,促进更谨慎的反思式推理。基于此,提出无需训练的自纠正框架 ESC(Emotional Self-Correction):通过外部验证器识别可能错误的初始回答,并注入情绪反馈以引导模型反思并生成更优修正结果。大量实验覆盖安全、幻觉、视觉感知和多模态推理等基准,结果表明 ESC 在不损失整体能力的前提下持续提升可靠性。这表明情绪不仅是可识别的能力,也可作为可扩展自纠正的实用控制信号。我们相信 ESC 为构建类人、情感融合的可靠模型研究提供了新方向。项目开源:https://genai4e.github.io/ESC/

原文摘要 · Abstract (English)

Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, yet they remain vulnerable to unreliable reasoning. Existing self-correction methods mitigate these issues but typically rely on post-training or carefully engineered feedback, incurring high computational cost. In this work, we revisit this challenge through the lens of emotional cues, asking whether they can activate latent self-correction behaviors in VLMs without additional training. \textbf{We find that emotional signals serve as an effective trigger for self-correction, encouraging more cautious and reflective reasoning}. Motivated by this finding, we propose \escabstract (\textbf{\underline{E}}motional \textbf{\underline{S}}elf-\textbf{\underline{C}}orrection), a training-free self-correction framework. ESC introduces an external verifier that detects potentially incorrect initial responses and injects emotional feedback to encourage model to reflect, and produce a better revised response without additional training. Extensive experiments across safety, hallucination, vision-centric perception, and multimodal reasoning benchmarks show that ESC consistently improves reliability while preserving overall model utility. These results suggest that emotion can function not only as an ability to be recognized, but also as a practical control signal for scalable self-correction in VLMs. \textbf{We therefore believe that ESC provides a strong foundation for a new reliable human-like, emotion-integrated research direction.} Our project is publicly available at \textcolor{red}{https://genai4e.github.io/ESC/}.

自纠正情绪建模视觉语言模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。