评估电影音频描述的主观性,提出新基准测试长片段理解与视觉欣赏能力。
What You See is What You Ask: Evaluating Audio Descriptions
- 构建基于问答的长视频段评测框架ADQA,测试几分钟级视频理解
- 发现现有自动生成模型远落后于人工描述,尤其在叙事连贯性上
- 适合无障碍技术、多模态生成研究者使用,推动真实场景评估
音频描述(AD)为视障和低视力(BLV)用户讲述电影中的重要视觉细节,帮助其理解剧情并欣赏画面。现有自动音频描述研究多聚焦于几秒的剪辑片段,且仅通过对比单一参考描述进行评估。然而,描述内容本身具有高度主观性。通过对同一部电影的两段独立音频描述进行对齐与分析,我们量化了描述时机、内容选择及重点突出的主观差异。结果表明,仅用剪辑片段评估不充分。为此,我们提出ADQA——一个针对数分钟连续视频段的问答评测基准,检验音频描述是否有助于BLV用户理解故事与欣赏视觉细节。ADQA包含关于视觉事实的视觉欣赏(VA)问题和基于情节的叙事理解(NU)问题。通过该基准,我们发现当前自动生成方法显著落后于人工描述。最后,我们提出未来研究建议,并开放公共排行榜用于持续评估。
原文摘要 · Abstract (English)
Audio descriptions (ADs) narrate important visual details in movies, enabling Blind and Low Vision (BLV) users to understand narratives and appreciate visual details. Existing works in automatic AD generation mostly focus on few-second trimmed clips, and evaluate them by comparing against a single ground-truth reference AD. However, writing ADs is inherently subjective. Through alignment and analysis of two independent AD tracks for the same movies, we quantify the subjectivity in when and whether to describe, and what and how to highlight. Thus, we show that working with trimmed clips is inadequate. We propose ADQA, a QA benchmark that evaluates ADs at the level of few-minute long, coherent video segments, testing whether they would help BLV users understand the story and appreciate visual details. ADQA features visual appreciation (VA) questions about visual facts and narrative understanding (NU) questions based on the plot. Through ADQA, we show that current AD generation methods lag far behind human-authored ADs. We conclude with several recommendations for future work and introduce a public leaderboard for benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。