首个仅靠视觉线索实现电影级自动配音的端到端系统
NowYouSee Me: Context-Aware Automatic Audio Description
- 基于时序增强与锚点检测,精准定位画面描述事件
- 自纠错模块使描述边界从粗到细逐步优化,提升准确性
- 无需依赖标注时间戳或元数据,适合无障碍内容生成
音频描述(AD)是保障多媒体内容可访问性的关键系统,通过在合适时机添加对视觉元素的解说,满足视障人群需求。本文提出首个统一的上下文感知自动音频描述系统 $[32m\mathrm{CA^3D}\u001b[0m$,可在长视频中精确定位并生成音频描述事件脚本。系统包含:1)时序特征增强模块,有效捕捉长期依赖;2)基于锚点的事件检测器与特征抑制模块,定位事件并提取生成所需特征;3)自修正模块,利用生成结果迭代优化事件边界。不同于依赖元数据或人工标注时间戳的传统方法,$[32m\mathrm{CA^3D}\u001b[0m$ 是首个仅使用视觉线索的端到端可训练系统。大量实验表明,该系统在事件检测与脚本生成任务上均超越现有架构,达成音频描述自动化新基准。
原文摘要 · Abstract (English)
Audio Description (AD) plays a pivotal role as an application system aimed at guaranteeing accessibility in multimedia content, which provides additional narrations at suitable intervals to describe visual elements, catering specifically to the needs of visually impaired audiences. In this paper, we introduce $\mathrm{CA^3D}$, the pioneering unified Context-Aware Automatic Audio Description system that provides AD event scripts with precise locations in the long cinematic content. Specifically, $\mathrm{CA^3D}$ system consists of: 1) a Temporal Feature Enhancement Module to efficiently capture longer term dependencies, 2) an anchor-based AD event detector with feature suppression module that localizes the AD events and extracts discriminative feature for AD generation, and 3) a self-refinement module that leverages the generated output to tweak AD event boundaries from coarse to fine. Unlike conventional methods which rely on metadata and ground truth AD timestamp for AD detection and generation tasks, the proposed $\mathrm{CA^3D}$ is the first end-to-end trainable system that only uses visual cue. Extensive experiments demonstrate that the proposed $\mathrm{CA^3D}$ improves existing architectures for both AD event detection and script generation metrics, establishing the new state-of-the-art performances in the AD automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。