arXiv:2506.15220cs.CVcs.CL2025-06被引 27

视频图文大模型新突破,生成更准更细的描述与问答。

video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models

  • 用多轮偏好优化动态更新参考策略,避免模型僵化。
  • 3B/7B模型在多个评测中达开源最优,72B超所有开源系统。
  • 适合做视频理解、图文生成与复杂问答任务的研究者使用。

我们提出 video-SALMONN 2,一个音频-视觉大语言模型系列,在视频描述与问答任务中达到新的开源最佳性能(SOTA)。核心贡献是多轮直接偏好优化(MrDPO),结合高质量字幕目标,同时奖励内容完整性和事实准确性。不同于传统DPO固定参考策略,MrDPO定期通过新初始化轻量适配器从最新偏好数据中重置参考,避免参考过时,实现持续优化。该方法生成的字幕比GPT-4o和Gemini-1.5 Pro等闭源系统更详细准确。我们进一步利用该模型构建高质量视频-字幕语料库,用于监督微调新模型,将优势扩展至复杂视频问答任务。在Video-MME、WorldSense、AVUT、Video-Holmes、DailyOmni、MLVU和LVBench等多个音视频与纯视觉理解基准上,3B和7B模型在可比规模下达到SOTA,72B模型超越所有其他开源系统。代码、模型与数据已公开于:https://github.com/bytedance/video-SALMONN-2。

原文摘要 · Abstract (English)

We present video-SALMONN 2, a family of audio-visual large language models that set new state-of-the-art (SOTA) results in video description and question answering (QA). Our core contribution is multi-round direct preference optimisation (MrDPO), paired with a caption-quality objective that jointly rewards completeness and factual accuracy. Unlike standard DPO with a fixed reference policy, MrDPO periodically refreshes the reference by bootstrapping from a newly re-initialised lightweight adapter trained on the latest preferences, avoiding reference staleness and enabling continual improvement. This strategy produces captions that are consistently more detailed and accurate than those from proprietary systems such as GPT-4o and Gemini-1.5 Pro. We further distil these gains by using our model to generate a high-quality video-caption corpus for supervised fine-tuning of new models, transferring benefits beyond captioning to strong performance on complex video-QA tasks. Across widely used audio-visual and visual-only understanding benchmarks (including Video-MME, WorldSense, AVUT, Video-Holmes, DailyOmni, MLVU, and LVBench), our 3B and 7B models achieve SOTA results at comparable scales, while the 72B model surpasses all other open-source systems. Our source code, models, and data are released at \href{https://github.com/bytedance/video-SALMONN-2}{https://github.com/bytedance/video-SALMONN-2}.

视频理解多模态大模型生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。