arXiv:2511.09448cs.MMcs.LG2025-11被引 1

让机器自动生成足球比赛音频描述,无需依赖人工标注。

MCAD: Multimodal Context-Aware Audio Description Generation For Soccer

  • 用电影音频描述数据微调视频大模型,学习叙事结构。
  • 结合球员身份、赛事动作和解说信息生成完整描述。
  • 新评估指标可量化分析描述质量,适合无障碍应用。

音频描述(AD)对视障人群获取视觉内容至关重要。现有方法多聚焦高质量电影,依赖人工标注的参考描述。本文提出端到端的MCAD系统,将音频描述扩展至足球赛事领域,无需依赖真实标注。为弥补领域专属数据缺失,我们基于公开电影AD数据集对视频大语言模型进行微调,使其掌握音频描述的叙述范式。推理时,MCAD融合球员身份、比赛事件与动作、赛事解说等多模态上下文信息,结合输入提示生成每段视频的完整音频描述。为此,我们设计新评估指标ARGE-AD,从五方面评估生成结果:(i)是否使用人名,(ii)是否提及动作与事件,(iii)长度是否合适,(iv)是否避免代词,(v)是否与解说或字幕重叠。我们在电影与足球数据集上进行了深入分析,并验证该指标在跨域评估中的有效性。此外,我们贡献了100段由两位专家标注的足球比赛音频描述。

原文摘要 · Abstract (English)

Audio Descriptions (AD) are essential for making visual content accessible to individuals with visual impairments. Recent works have shown a promising step towards automating AD, but they have been limited to describing high-quality movie content using human-annotated ground truth AD in the process. In this work, we present an end-to-end pipeline, MCAD, that extends AD generation beyond movies to the domain of sports, with a focus on soccer games, without relying on ground truth AD. To address the absence of domain-specific AD datasets, we fine-tune a Video Large Language Model on publicly available movie AD datasets so that it learns the narrative structure and conventions of AD. During inference, MCAD incorporates multimodal contextual cues such as player identities, soccer events and actions, and commentary from the game. These cues, combined with input prompts to the fine-tuned VideoLLM, allow the system to produce complete AD text for each video segment. We further introduce a new evaluation metric, ARGE-AD, designed to accurately assess the quality of generated AD. ARGE-AD evaluates the generated AD for the presence of five characteristics: (i) usage of people's names, (ii) mention of actions and events, (iii) appropriate length of AD, (iv) absence of pronouns, and (v) overlap from commentary or subtitles. We present an in-depth analysis of our approach on both movie and soccer datasets. We also validate the use of this metric to quantitatively comment on the quality of generated AD using our metric across domains. Additionally, we contribute audio descriptions for 100 soccer game clips annotated by two AD experts.

音频描述足球多模态视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。