arXiv:2412.17562cs.CL2024-12被引 1

构建7.5万句英-罗马乌尔都语平行数据集,助力低资源语言处理

ERUPD -- English to Roman Urdu Parallel Dataset

  • 混合生成与真实对话数据,提升罗马乌尔都语表达多样性
  • 通过人工评估修正语音、混用词和同义表达不一致问题
  • 适合机器翻译、情感分析及多语言教育研究者使用

弥合语言鸿沟有助于全球发展与文化交流。本文针对罗马乌尔都语——一种广泛用于数字交流的拉丁字母拼写乌尔都语——因缺乏标准化、语音变异及与英语混用带来的处理难题,构建了一个包含75,146句对的新颖平行数据集。研究采用混合方法,结合高级提示工程生成的合成数据与个人聊天群组的真实对话数据,并通过人工评估阶段修正语言不一致问题,确保代码混用、音标表征和同义表达的准确性。该数据集充分捕捉了罗马乌尔都语的多样语言特征,可为机器翻译、情感分析和多语言教育提供关键支持。

原文摘要 · Abstract (English)

Bridging linguistic gaps fosters global growth and cultural exchange. This study addresses the challenges of Roman Urdu -- a Latin-script adaptation of Urdu widely used in digital communication -- by creating a novel parallel dataset comprising 75,146 sentence pairs. Roman Urdu's lack of standardization, phonetic variability, and code-switching with English complicates language processing. We tackled this by employing a hybrid approach that combines synthetic data generated via advanced prompt engineering with real-world conversational data from personal messaging groups. We further refined the dataset through a human evaluation phase, addressing linguistic inconsistencies and ensuring accuracy in code-switching, phonetic representations, and synonym variability. The resulting dataset captures Roman Urdu's diverse linguistic features and serves as a critical resource for machine translation, sentiment analysis, and multilingual education.

机器翻译低资源语言数据集罗马乌尔都语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。