不训练模型,用镜头语法生成电影音频描述,效果超精调方法。
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
- 以镜头为单位,融合镜头尺度和叙事结构指导描述生成。
- 在多个基准上超越此前无训练方法,部分超过微调模型表现。
- 提出新评估指标与多候选生成协议,更贴近实际使用场景。
本文旨在为剪辑过的视频内容(如电影、电视剧)自动生成音频描述(ADs)。为此,我们提出一种两阶段框架,以“镜头”作为视频理解的基本单元,扩展时间上下文至邻近镜头,并引入镜头尺度、叙事结构等电影语法机制来引导描述生成。该方法兼容开源与专有视觉语言模型(VLMs),通过附加模块融入领域知识,无需对VLM进行额外训练。在多个基准测试中,其性能优于所有已有无训练方法,甚至在部分数据集上超越微调模型。为评估生成描述的质量,我们提出一种针对动作描述的“动作得分”新指标;同时设计了一种新评估协议,将自动系统视为音频描述助手,要求生成多个候选描述供选择,以更真实反映应用需求。
原文摘要 · Abstract (English)
Our objective is the automatic generation of Audio Descriptions (ADs) for edited video material, such as movies and TV series. To achieve this, we propose a two-stage framework that leverages "shots" as the fundamental units of video understanding. This includes extending temporal context to neighbouring shots and incorporating film grammar devices, such as shot scales and thread structures, to guide AD generation. Our method is compatible with both open-source and proprietary Visual-Language Models (VLMs), integrating expert knowledge from add-on modules without requiring additional training of the VLMs. We achieve state-of-the-art performance among all prior training-free approaches and even surpass fine-tuned methods on several benchmarks. To evaluate the quality of predicted ADs, we introduce a new evaluation measure -- an action score -- specifically targeted to assessing this important aspect of AD. Additionally, we propose a novel evaluation protocol that treats automatic frameworks as AD generation assistants and asks them to generate multiple candidate ADs for selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。