arXiv:2507.14543cs.CVcs.CY2025-07

实时将视频会议中的手语翻译为字幕,助力听障者沟通

Real Time Captioning of Sign Language Gestures in Video Meetings

  • 基于浏览器插件实现实时手语转字幕
  • 使用超过2000个词级美式手语视频数据集训练
  • 适合视频会议中听障用户及需无障碍沟通的场景

与听力障碍者沟通始终是一项挑战,手语是主要沟通方式之一。然而,多数人并不熟悉手语的细微差别。计算机视觉驱动的手语识别旨在消除聋哑人与普通人之间的沟通障碍。疫情期间,视频会议成为主流沟通方式,研究发现听障人士更倾向于在会议中用手语而非打字。本文提出一款浏览器扩展,可自动将视频会议中的手语实时翻译为字幕,供其他参会者查看。系统基于包含2000多个词级美式手语(ASL)视频的大规模数据集,由100多名手语者录制,实现高精度实时转换。

原文摘要 · Abstract (English)

It has always been a rather tough task to communicate with someone possessing a hearing impairment. One of the most tested ways to establish such a communication is through the use of sign based languages. However, not many people are aware of the smaller intricacies involved with sign language. Sign language recognition using computer vision aims at eliminating the communication barrier between deaf-mute and ordinary people so that they can properly communicate with others. Recently the pandemic has left the whole world shaken up and has transformed the way we communicate. Video meetings have become essential for everyone, even people with a hearing disability. In recent studies, it has been found that people with hearing disabilities prefer to sign over typing during these video calls. In this paper, we are proposing a browser extension that will automatically translate sign language to subtitles for everyone else in the video call. The Large-scale dataset which contains more than 2000 Word-Level ASL videos, which were performed by over 100 signers will be used.

手语识别实时翻译视频会议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。