首个支持视频推理的开源多模态大模型,提升复杂视频理解能力。
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
- 构建步骤式推理数据集与对比步选择优化方法,实现细粒度奖励建模
- 在4000+专家标注问题上,比基线模型高3-8%准确率
- 零样本检测合成视频,适合视频内容分析与安全应用
现有推理优化主要集中在数学问题和视觉图像输入,忽视了通用视频理解的应用。本文提出 video-SALMONN-o1,首个面向通用视频理解任务的开源推理增强型音视频大语言模型。为提升推理能力,我们构建了一个包含挑战性音视频问题及分步解答的推理密集型数据集,并提出过程直接偏好优化(pDPO),通过对比步选择实现针对多模态输入的高效步级奖励建模。此外,我们引入 RivaBench,首个推理密集型视频理解基准,涵盖超过4,000个高质量、专家标注的问答对,覆盖单口喜剧、学术演讲和合成视频检测等场景。video-SALMONN-o1 在多个视频推理基准上相较 LLaVA-OneVision 基线模型提升3-8%准确率;pDPO 在 RivaBench 上相较监督微调模型提升6-8%。增强的推理能力使 video-SALMONN-o1 具备零样本合成视频检测能力。
原文摘要 · Abstract (English)
While recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in general video understanding.This paper proposes video-SALMONN-o1, the first open-source reasoning-enhanced audio-visual LLM designed for general video understanding tasks. To enhance its reasoning abilities, we develop a reasoning-intensive dataset featuring challenging audio-visual questions with step-by-step solutions. We also propose process direct preference optimization (pDPO), which leverages contrastive step selection to achieve efficient step-level reward modelling tailored for multimodal inputs. Additionally, we introduce RivaBench, the first reasoning-intensive video understanding benchmark, featuring over 4,000 high-quality, expert-curated question-answer pairs across scenarios such as standup comedy, academic presentations, and synthetic video detection. video-SALMONN-o1 achieves 3-8% accuracy improvements over the LLaVA-OneVision baseline across different video reasoning benchmarks. Besides, pDPO achieves 6-8% improvements compared to the supervised fine-tuning model on RivaBench. Enhanced reasoning enables video-SALMONN-o1 zero-shot synthetic video detection capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。