arXiv:2510.10395cs.CV2025-10被引 25

让音视频描述更精准对齐,提升理解与生成效果

AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

  • 通过音视频时间协同机制,优化跨模态对齐
  • 在4个基准上超越现有开源模型,最高提升12.3%
  • 适合需要精确时序描述的视频理解与生成任务

音视频视频字幕生成旨在生成语义丰富且视听事件时间对齐的描述,有助于视频理解与生成。本文提出AVoCaDO,一种由音视频模态间时间协同驱动的强音视频字幕生成模型。我们设计了两阶段后训练流程:(1) AVoCaDO SFT,基于新构建的107,000条高质量、时间对齐的音视频字幕数据集进行微调;(2) AVoCaDO GRPO,利用定制奖励函数进一步增强时间连贯性与对话准确性,同时正则化字幕长度并减少生成坍缩。实验表明,AVoCaDO在四个音视频字幕基准上显著优于现有开源模型,在仅视觉设置下的VDC与DREAM-1K基准上也达到具有竞争力的表现。

原文摘要 · Abstract (English)

Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a powerful audiovisual video captioner driven by the temporal orchestration between audio and visual modalities. We propose a two-stage post-training pipeline: (1) AVoCaDO SFT, which fine-tunes the model on a newly curated dataset of 107K high-quality, temporally-aligned audiovisual captions; and (2) AVoCaDO GRPO, which leverages tailored reward functions to further enhance temporal coherence and dialogue accuracy while regularizing caption length and reducing collapse. Experimental results demonstrate that AVoCaDO significantly outperforms existing open-source models across four audiovisual video captioning benchmarks, and also achieves competitive performance on the VDC and DREAM-1K benchmark under visual-only settings.

音视频理解字幕生成时间对齐多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。