用双层次注意力提升长视频语音描述连贯性
DANTE-AD: Dual-Vision Attention Network for Long-Term Audio Description
- 融合帧级与场景级特征,构建双视角注意力机制
- 在电影片段上实现超越现有方法的描述准确率
- 适合需要长期上下文理解的无障碍视频解说场景
音频描述是为视障观众提供视频关键视觉元素解说的叙述性内容。尽管短时视频理解已快速发展,但保持长期视觉叙事连贯性的解决方案仍待解决。现有方法仅依赖帧级嵌入,虽能描述物体信息,却缺乏跨场景上下文。我们提出DANTE-AD,一种基于双视觉Transformer架构的增强型视频描述模型,通过顺序融合帧级与场景级嵌入,提升长期上下文理解能力。提出一种新颖的序列交叉注意力机制,实现细粒度音频描述生成的上下文定位。在多个知名电影片段的关键场景上评估,DANTE-AD在传统NLP指标和基于大语言模型的评估中均优于现有方法。
原文摘要 · Abstract (English)
Audio Description is a narrated commentary designed to aid vision-impaired audiences in perceiving key visual elements in a video. While short-form video understanding has advanced rapidly, a solution for maintaining coherent long-term visual storytelling remains unresolved. Existing methods rely solely on frame-level embeddings, effectively describing object-based content but lacking contextual information across scenes. We introduce DANTE-AD, an enhanced video description model leveraging a dual-vision Transformer-based architecture to address this gap. DANTE-AD sequentially fuses both frame and scene level embeddings to improve long-term contextual understanding. We propose a novel, state-of-the-art method for sequential cross-attention to achieve contextual grounding for fine-grained audio description generation. Evaluated on a broad range of key scenes from well-known movie clips, DANTE-AD outperforms existing methods across traditional NLP metrics and LLM-based evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。