arXiv:2502.20480cs.CVcs.HC2025-02被引 22

用大模型生成更适配视障用户的视频描述,数据集达4万条。

VideoA11y: Method and Dataset for Accessible Video Description

  • 结合多模态大模型与无障碍指南生成描述
  • 4万条数据集使描述质量媲美专业人类标注
  • 适合无障碍研究、AI辅助视觉设计者使用

视频描述对盲人和低视力(BLV)用户获取视觉内容至关重要。然而,当前人工智能模型生成的描述常因训练数据中人工标注质量不足而无法满足需求。为此,我们提出 VideoA11y 方法,利用多模态大语言模型(MLLMs)和视频无障碍指南,生成专为 BLV 用户定制的描述。基于此方法,我们构建了规模最大的 40,000 条视频描述数据集——VideoA11y-40K。在 15 个视频类别中,通过 347 名视力正常参与者、40 名 BLV 参与者及 7 名专业描述员的严格实验表明,VideoA11y 的描述在清晰度、准确性、客观性、描述性和用户满意度方面均优于新手人类标注,并达到与受训人类标注相当的水平。我们在 VideoA11y-40K 上使用标准与自定义指标评估模型,证明在该数据集上微调的 MLLMs 能生成高质量的可访问描述。代码与数据集已公开:https://people-robots.github.io/VideoA11y。

原文摘要 · Abstract (English)

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human annotations within training datasets, resulting in descriptions that do not fully meet BLV users' needs. To address this gap, we introduce VideoA11y, an approach that leverages multimodal large language models (MLLMs) and video accessibility guidelines to generate descriptions tailored for BLV individuals. Using this method, we have curated VideoA11y-40K, the largest and most comprehensive dataset of 40,000 videos described for BLV users. Rigorous experiments across 15 video categories, involving 347 sighted participants, 40 BLV participants, and seven professional describers, showed that VideoA11y descriptions outperform novice human annotations and are comparable to trained human annotations in clarity, accuracy, objectivity, descriptiveness, and user satisfaction. We evaluated models on VideoA11y-40K using both standard and custom metrics, demonstrating that MLLMs fine-tuned on this dataset produce high-quality accessible descriptions. Code and dataset are available at https://people-robots.github.io/VideoA11y.

视频描述无障碍多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。