让视频字幕同时讲清画面、语音和声音的互动关系。
AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

- 通过显式绑定视听事件,实现多模态协同描述
- 在非语音声音覆盖上提升显著,跨模态关联更准确
- 适合需要精细音画联动理解的场景,如视频内容分析
全模态视频字幕不仅需融合视觉描述与语音转录,更需揭示视觉动作、语言、音乐与音效如何协同演进。现有大模型常将音视频视为松散关联,依赖语音识别,忽视非语音声音及其与视觉事件的联系。本文提出AVSCap框架:首先构建包含13万条三模态数据的训练集(AVSCap-130K),通过解耦再融合流程,先锚定视听证据再生成有依据的全模态字幕;其次训练70亿参数的字幕生成器(AVSCap-7B),采用两阶段策略——监督微调建立基础能力,样本高效强化学习结合混合奖励优化音频完整性和音画协同性;最后提出AVSCapBench基准,将字幕分解为视觉、音频与协同事件,以细粒度事件召回率评估。在该基准及外部评测中,AVSCap-7B在非语音声音覆盖率与跨模态绑定上表现最优,是当前开源模型中的最佳性能者。
原文摘要 · Abstract (English)
Omni-modal video captioning is not merely combining visual captioning with audio transcription: a useful caption must describe how visual actions, speech, music, and sound effects co-evolve. Existing large multimodal models often fail at this relational step, treating audio and visual streams as loosely coupled observations, relying on automatic speech recognition, and under-specifying non-speech sounds and their links to visual events. We present AVSCap, a framework for audio-visual captioning centered on explicit cross-modal event binding. First, we construct AVSCap-130K, a tri-modal training corpus generated by a decoupled-then-fused pipeline that anchors visual and acoustic evidence before composing grounded omni-modal captions. Second, we train AVSCap-7B, a 7B captioner with a two-stage strategy: supervised fine-tuning establishes baseline capabilities, while sample-efficient reinforcement learning uses hybrid rewards to optimize acoustic completeness and audio-visual synergy. Our scaling analysis shows that reinforcement learning brings larger gains than increasing SFT data. Third, we introduce AVSCapBench, a benchmark that decomposes captions into visual, audio, and synergy events and evaluates them with fine-grained event recall. Experiments on AVSCapBench and external benchmarks show that AVSCap-7B improves non-speech audio coverage and cross-modal binding, delivering the best overall performance among evaluated open-source models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。