arXiv:2504.02293cs.CLcs.AI2025-04被引 6

首个孟加拉语手语翻译数据集,助力聋哑人群语言沟通

Breaking the Silence: A Dataset and Benchmark for Bangla Text-to-Gloss Translation

  • 构建1000条人工标注+4000条合成的孟加拉语-手语词对数据集
  • 合成数据有效提升低资源手语翻译性能,模型准确率显著提高
  • 适合关注残障包容、低资源语言处理的研究者与开发者

手语是手语与对应口语之间的书面近似。尽管孟加拉国聋哑和听力障碍人口至少达300万,但孟加拉手语(BdSL)仍研究不足,此前无相关文本到手语词翻译工作,也无公开数据集。为填补这一空白,我们构建了首个孟加拉语文本到手语词的数据集,包含1000条人工标注和4000条合成生成的句子-手语词对,以及159条专家标注的测试集。实验对比多个微调开源模型与领先闭源大模型在低资源BdSL翻译中的表现。GPT-5.4整体表现最佳,而微调后的mBART模型虽体积约小100%,仍具竞争力。Qwen-3在人类评估中优于所有模型。本工作首次提出孟加拉语文本到手语词翻译的数据集与训练模型,并证明系统性合成数据在解决低资源手语翻译挑战中的有效性。

原文摘要 · Abstract (English)

Gloss is a written approximation that bridges Sign Language (SL) and its corresponding spoken language. Despite a deaf and hard-of-hearing population of at least 3 million in Bangladesh, Bangla Sign Language (BdSL) remains largely understudied, with no prior work on Bangla text-to-gloss translation and no publicly available datasets. To address this gap, we construct the first Bangla text-to-gloss dataset, consisting of 1,000 manually annotated and 4,000 synthetically generated Bangla sentence-gloss pairs, along with 159 expert human-annotated pairs used as a test set. Our experimental framework performs a comparative analysis between several fine-tuned open-source models and a leading closed-source LLM to evaluate their performance in low-resource BdSL translation. GPT-5.4 achieves the best overall performance, while a fine-tuned mBART model performs competitively despite being approximately 100% smaller. Qwen-3 outperforms all other models in human evaluation. This work introduces the first dataset and trained model for Bangla text-to-gloss translation. It also demonstrates the effectiveness of systematically generated synthetic data for addressing challenges in low-resource sign language translation.

手语翻译低资源语言数据集孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。