arXiv:2602.01390cs.HCcs.AI2026-02中稿 · ASSETS 2026

用心理测量学方法评估音描述质量,实现大规模高效质检。

Toward Scalable Audio Description Quality Control: A Workflow for Evaluating Human and VLM Raters

  • 基于项目反应理论构建多维度评估框架,量化人与AI评分员能力。
  • 顶尖VLM表现接近人类,但推理可靠性仍逊于真人。
  • 适合无障碍内容审核、AI辅助质检系统开发者参考。

数字视频在通信、教育和娱乐中至关重要,但缺乏音频描述(AD)会使视障用户被排除在外。尽管众包平台和视觉语言模型(VLMs)扩展了AD生成能力,但质量评估常缺乏系统性。现有方法依赖NLP指标和短片段指南,无法有效评估长时序AD质量。为此,我们提出一种基于项目反应理论的方法论流程,以专家建立的基准为参照,评估人类与VLM评分员的水平。评估基于六维框架,依据专业指南并融合无障碍专家与视障顾问的洞察。结果表明,表现最佳的VLM可达到与人类评分员相当的基准评分水平。然而定性分析显示,VLM的推理过程较人类更不可靠且难以行动。这些发现凸显了融合VLM与人工监督的混合评估系统的潜力,为实现可扩展的音频描述质量控制提供了可行路径。

原文摘要 · Abstract (English)

Digital video is central to communication, education, and entertainment, but without audio description (AD), blind and low-vision users are excluded. While crowdsourced platforms and vision-language models (VLMs) expand AD production, quality is rarely checked systematically. Existing evaluations rely on NLP metrics and short-clip guidelines, leaving open the question of how to assess long-form AD quality at scale. To address this, we developed a methodological workflow using Item Response Theory to evaluate VLM and human rater proficiency against expert-established ground truth. Evaluations were based on a six-dimensional framework, grounded in professional guidelines and shaped by insights from our accessibility experts and blind consultants. Findings suggest that top-performing VLMs can approximate ground-truth ratings at levels comparable to human raters. However, qualitative analysis reveals that VLM reasoning is less reliable and actionable than that of human respondents. These insights underscore the potential of hybrid evaluation systems that leverage VLMs alongside human oversight, offering a path toward scalable AD quality control.

音频描述质量评估VLM可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。