首个印度语音频描述数据集,助力视障群体无障碍观影。
Andha-Dhun: A First Look at Audio Descriptions in Hindi

- 构建首个8部电影的原创印地语音频描述数据集
- 直接翻译英文描述易丢失文化细节,人工翻译更优但仍有不足
- 强调适配本地受众比严格对齐原文更重要
音频描述(ADs)在音视频媒体间隙为盲人和低视力(BLV)观众提供视觉内容解说。随着印度中央电影认证委员会(CBFC)推动政策落地,需将ADs从英语扩展至印地语等本土语言。然而此前尚无针对印度语言的生成研究。本文首次系统开展印地语音频描述研究,提出Andha-Dhun数据集,收录8部全长电影的真人撰写的印地语音频描述。探索两种生成方式:(i)由英文密集视频描述直接生成,(ii)将英文音频描述翻译为印地语。采用困惑度与大模型评判指标分别评估流畅性与质量。对比分析同时具备英、印地语真人描述的影片发现,直接翻译会产生伪影并降低多样性;机器翻译难以适应文化语境,而人工翻译虽更优但仍不理想。结果表明,印地语音频描述的核心目标是服务印度视障人群,需以受众适配优先于源文本忠实度。
原文摘要 · Abstract (English)
Audio Descriptions (ADs) narrate visual content for Blind and Low Vision (BLV) audiences during gaps in audiovisual media. There is growing momentum around ADs in movies and TV shows, and with mandates from India's Central Board of Film Certification (CBFC), there is a need to expand ADs beyond English. Yet, there is no work that generates ADs for any Indian language. To address this gap, we present the first systematic study of ADs in Hindi, contributing to aspects such as data, generation, and evaluation. We introduce Andha-Dhun, the first dataset of human-authored Hindi ADs collected from 8 full-length movies. We explore two approaches for generating ADs in Hindi: (i) directly from English dense video descriptions, and (ii) translating English ADs into Hindi. We evaluate these approaches using perplexity and LLM-as-a-judge metrics to assess fluency and quality respectively. We also analyze movies that have both English and Hindi human-authored ADs and find that naive translation introduces artifacts and narrows diversity compared to original Hindi ADs. Direct machine translation fails to adapt cultural references, while human-translated ADs do better but still fall short. Our findings emphasize that the purpose of Hindi ADs is accessibility for Indian BLV audiences, and that this requires adapting content for the audience more than strict fidelity to the source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。