arXiv:2511.08507cs.CLcs.AI2025-11被引 1

构建首个孟加拉手语句级翻译数据集,助力低资源手语研究。

Introducing A Bangla Sentence - Gloss Pair Dataset for Bangla Sign Language Translation and Research

  • 基于规则的检索增强生成,合成3000+句对扩充数据集
  • 1000条人工标注句-词序列对,由专业手语者验证质量
  • 首次在真实数据上对比多种模型的手语翻译一致性

孟加拉手语(BdSL)翻译因缺乏大规模句级数据而成为低资源NLP任务,现有研究多限于单词与字母层面。本文提出Bangla-SGP,一个包含1000条人工标注的句-词序列对的新平行数据集,并通过基于语法和形态规则的规则化检索增强生成(RAG)管道,合成约3000条补充数据。词序列由支持孟加拉手语的单个词构成,作为连续手语的中间表示。数据集包含1000条高质量孟加拉语句子,由专业手语者手动标注为词序列。增强过程结合了语言学策略与提示工程,基于对人工标注数据的分析及与专业手语者的密切合作。我们微调了mBart50、Google mT5、GPT4.1-nano等多个Transformer模型,并基于BLEU分数评估其句到词序列的翻译性能,比较模型在本数据集与RWTH-PHOENIX-2014T基准上的词序列一致性。

原文摘要 · Abstract (English)

Bangla Sign Language (BdSL) translation represents a low-resource NLP task due to the lack of large-scale datasets that address sentence-level translation. Correspondingly, existing research in this field has been limited to word and alphabet level detection. In this work, we introduce Bangla-SGP, a novel parallel dataset consisting of 1,000 human-annotated sentence-gloss pairs which was augmented with around 3,000 synthetically generated pairs using syntactic and morphological rules through a rule-based Retrieval-Augmented Generation (RAG) pipeline. The gloss sequences of the spoken Bangla sentences are made up of individual glosses which are Bangla sign supported words and serve as an intermediate representation for a continuous sign. Our dataset consists of 1000 high quality Bangla sentences that are manually annotated into a gloss sequence by a professional signer. The augmentation process incorporates rule-based linguistic strategies and prompt engineering techniques that we have adopted by critically analyzing our human annotated sentence-gloss pairs and by working closely with our professional signer. Furthermore, we fine-tune several transformer-based models such as mBart50, Google mT5, GPT4.1-nano and evaluate their sentence-to-gloss translation performance using BLEU scores, based on these evaluation metrics we compare the model's gloss-translation consistency across our dataset and the RWTH-PHOENIX-2014T benchmark.

手语翻译低资源数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。