用幻觉纠正训练模型,提升视频与文本对齐能力。
Can Hallucination Correction Improve Video-Language Alignment?
- 自训练框架自动识别并修正描述与视频不符的幻觉
- 在视频-字幕绑定和文本到视频检索任务中表现更优
- 适合研究视觉语言对齐与生成质量优化的学者
大型视觉语言模型常生成与视觉输入不符的幻觉内容。以往工作聚焦于减少幻觉,本文则探索将幻觉纠正作为训练目标以增强视频与语言对齐。提出HACA自训练框架,学习修正与视频内容不一致的描述。通过识别并修正不一致性,HACA提升了模型在时空推理中对视频与文本表征的对齐能力。实验结果表明,在视频-字幕绑定和文本到视频检索任务中均取得稳定提升,证明基于幻觉纠正的任务是改善视觉与语言对齐的有效策略。
原文摘要 · Abstract (English)
Large Vision-Language Models often generate hallucinated content that is not grounded in its visual inputs. While prior work focuses on mitigating hallucinations, we instead explore leveraging hallucination correction as a training objective to improve video-language alignment. We introduce HACA, a self-training framework learning to correct hallucinations in descriptions that do not align with the video content. By identifying and correcting inconsistencies, HACA enhances the model's ability to align video and textual representations for spatio-temporal reasoning. Our experimental results show consistent gains in video-caption binding and text-to-video retrieval tasks, demonstrating that hallucination correction-inspired tasks serve as an effective strategy for improving vision and language alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。