首个支持瑞士四种语言的音频描述翻译系统,提升视障人群信息获取便利性。
SwissADT: An Audio Description Translation System for Swiss Languages
- 融合文本与视频信息,用大模型实现多语言音频描述自动翻译。
- 在德、法、意、英四语间测试,人工与自动评估均表现良好。
- 适合需要跨语言无障碍服务的研究者与公共机构参考。
音频描述(AD)是为视障人士提供的重要无障碍服务,通过声音传递视觉信息。尽管多语言机器翻译研究取得进展,但缺乏高质量且时间同步的音频描述数据,制约了瑞士等多语言国家的音频描述翻译(ADT)系统发展。此外,现有系统多依赖纯文本,尚未明确视频视觉信息是否能提升翻译质量。本文提出瑞士首个面向德语、法语、意大利语和英语的ADT系统SwissADT,通过收集带有视频片段的高质量音频描述数据,并结合大语言模型(LLMs)能力,实现音频描述脚本在瑞士主要语言间的自动翻译。大量实验结果(包括自动评估与人工评估)表明,SwissADT在多语言环境下具备显著潜力。我们认为,结合人类专家经验与大模型生成能力,可进一步提升系统性能,惠及更广泛的多语言用户群体。
原文摘要 · Abstract (English)
Audio description (AD) is a crucial accessibility service provided to blind persons and persons with visual impairment, designed to convey visual information in acoustic form. Despite recent advancements in multilingual machine translation research, the lack of well-crafted and time-synchronized AD data impedes the development of audio description translation (ADT) systems that address the needs of multilingual countries such as Switzerland. Furthermore, since the majority of ADT systems rely solely on text, uncertainty exists as to whether incorporating visual information from the corresponding video clips can enhance the quality of ADT outputs. In this work, we present SwissADT, the first ADT system implemented for three main Swiss languages and English. By collecting well-crafted AD data augmented with video clips in German, French, Italian, and English, and leveraging the power of Large Language Models (LLMs), we aim to enhance information accessibility for diverse language populations in Switzerland by automatically translating AD scripts to the desired Swiss language. Our extensive experimental ADT results, composed of both automatic and human evaluations of ADT quality, demonstrate the promising capability of SwissADT for the ADT task. We believe that combining human expertise with the generation power of LLMs can further enhance the performance of ADT systems, ultimately benefiting a larger multilingual target population.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。