构建首个高精度孟加拉语-英语句法时态标注语料库,助力低资源语言研究。
BiST: A Gold Standard Bangla-English Bilingual Corpus for Sentence Structure and Tense Classification with Inter-Annotator Agreement
- 构建30,534句双语语料,涵盖句法结构与时态两类标注
- 标注一致性达κ=0.82(句法)和κ=0.88(时态),质量可靠
- 适用于语法建模、自动反馈生成等实际任务,适合多语言研究者
高质量双语资源是推动低资源语言多语言自然语言处理的关键瓶颈,尤其对孟加拉语而言。为缓解这一缺口,我们推出BiST——一个严格构建的孟加拉语-英语语料库,用于句子级语法分类,涵盖两个核心维度:句法结构(简单、复杂、并列、复杂并列)与时态(现在、过去、未来)。语料来自开放许可的百科资源与自然对话文本,经系统预处理与自动语言识别,共包含30,534句,其中17,465句为英文,13,069句为孟加拉语。通过三名独立标注者与分维度的Fleiss Kappa(κ)评估,确保标注质量,句法与时态标注κ值分别为0.82与0.88。统计分析显示结构与时态分布合理;基线实验表明,融合语言特异性表示的双编码器架构优于强健的多语言编码器。该语料不仅可用于基准测试,还提供明确的语言学监督,支持语法建模任务,如可控文本生成、自动反馈生成与跨语言表征学习。BiST建立统一双语语法建模资源,推动基于语言学的多语言研究。
原文摘要 · Abstract (English)
High-quality bilingual resources remain a critical bottleneck for advancing multilingual NLP in low-resource settings, particularly for Bangla. To mitigate this gap, we introduce BiST, a rigorously curated Bangla-English corpus for sentence-level grammatical classification, annotated across two fundamental dimensions: syntactic structure (Simple, Complex, Compound, Complex-Compound) and tense (Present, Past, Future). The corpus is compiled from open-licensed encyclopedic sources and naturally composed conversational text, followed by systematic preprocessing and automated language identification, resulting in 30,534 sentences, including 17,465 English and 13,069 Bangla instances. Annotation quality is ensured through a multi-stage framework with three independent annotators and dimension-wise Fleiss Kappa ($κ$) agreement, yielding reliable and reproducible labels with $κ$ values of 0.82 and 0.88 for structural and temporal annotation, respectively. Statistical analyses demonstrate realistic structural and temporal distributions, while baseline evaluations show that dual-encoder architectures leveraging complementary language-specific representations consistently outperform strong multilingual encoders. Beyond benchmarking, BiST provides explicit linguistic supervision that supports grammatical modeling tasks, including controlled text generation, automated feedback generation, and cross-lingual representation learning. The corpus establishes a unified resource for bilingual grammatical modeling and facilitates linguistically grounded multilingual research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。