arXiv:2508.20655cs.CVcs.CL2025-08EMNLP被引 3

让视觉语言模型自我纠错,减少幻觉并提升安全性。

Improving Alignment in LVLMs with Debiased Self-Judgment

  • 模型内部生成去偏自评分数,无需外部数据。
  • 在多个基准上幻觉率降低,安全性和能力显著提升。
  • 适合需要低成本高安全性的多模态应用开发者。

大型语言模型(LLMs)和大型视觉-语言模型(LVLMs)的快速发展为视觉与语言模态的融合开辟了新机遇。然而,有效对齐这些模态仍具挑战性,常导致生成结果脱离视觉输入(即幻觉),引发各领域的安全问题。现有对齐方法如指令微调和偏好微调通常依赖外部数据集、人工标注或复杂后处理,限制可扩展性并增加成本。为此,我们提出一种新方法,通过模型内部生成去偏自评分数(debiased self-judgment score),实现无需外部资源的自主对齐优化。该机制改进了解码策略与偏好微调过程,在减少幻觉、增强安全性和整体性能方面表现优异。实证结果表明,本方法显著优于传统方法,为LVLM对齐提供了更高效可行的解决方案。

原文摘要 · Abstract (English)

The rapid advancements in Large Language Models (LLMs) and Large Visual-Language Models (LVLMs) have opened up new opportunities for integrating visual and linguistic modalities. However, effectively aligning these modalities remains challenging, often leading to hallucinations--where generated outputs are not grounded in the visual input--and raising safety concerns across various domains. Existing alignment methods, such as instruction tuning and preference tuning, often rely on external datasets, human annotations, or complex post-processing, which limit scalability and increase costs. To address these challenges, we propose a novel approach that generates the debiased self-judgment score, a self-evaluation metric created internally by the model without relying on external resources. This enables the model to autonomously improve alignment. Our method enhances both decoding strategies and preference tuning processes, resulting in reduced hallucinations, enhanced safety, and improved overall capability. Empirical results show that our approach significantly outperforms traditional methods, offering a more effective solution for aligning LVLMs.

视觉语言模型对齐优化自评机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。