arXiv:2504.05810cs.CV2025-04被引 15

通过提示感知的多实例学习,减少视频生成中的幻觉问题。

PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning

  • 基于提示上下文选择增强片段,动态生成拒收样本
  • 仅用1万条SFT数据提升5.3%效果,超越GPT-4o
  • 无需额外参数或人工标注,适配现有视频大模型

直接偏好优化(DPO)可降低视频多模态大模型(VLLM)的幻觉,但依赖离线偏好数据导致适应性差,且难以捕捉真实视频-回复错位。我们提出在线偏好学习框架VDPO,利用视频增强生成拒收样本,保持响应固定。为避免无效增强引发误拒,引入提示感知多实例学习的PaMi-VDPO:基于提示上下文筛选增强片段,构建候选集并采用近到远策略,先确保语义相关,再优先选择最符合提示的差异片段。该方法有效捕捉有意义视觉差异,抑制幻觉,避免误拒,提升对齐性。PaMi-VDPO无缝集成至现有VLLM,无需额外参数或人工标注。仅使用10,000条SFT数据,在VideoHallucer上比基础模型提升5.3%,超过GPT-4o,同时在通用视频基准测试中保持稳定表现。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) helps reduce hallucinations in Video Multimodal Large Language Models (VLLMs), but its reliance on offline preference data limits adaptability and fails to capture true video-response misalignment. We propose Video Direct Preference Optimization (VDPO), an online preference learning framework that eliminates the need for preference annotation by leveraging video augmentations to generate rejected samples while keeping responses fixed. However, selecting effective augmentations is non-trivial, as some clips may be semantically identical to the original under specific prompts, leading to false rejections and disrupting alignment. To address this, we introduce Prompt-aware Multi-instance Learning VDPO (PaMi-VDPO), which selects augmentations based on prompt context. Instead of a single rejection, we construct a candidate set of augmented clips and apply a close-to-far selection strategy, initially ensuring all clips are semantically relevant while then prioritizing the most prompt-aware distinct clip. This allows the model to better capture meaningful visual differences, mitigating hallucinations, while avoiding false rejections, and improving alignment. PaMi-VDPOseamlessly integrates into existing VLLMs without additional parameters, GPT-4/human supervision. With only 10k SFT data, it improves the base model by 5.3% on VideoHallucer, surpassing GPT-4o, while maintaining stable performance on general video benchmarks.

视频生成幻觉抑制偏好学习VLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。