arXiv:2508.05722cs.CL2025-08被引 2

构建首个医疗领域英阿平行语料库,助力精准翻译与跨语言研究。

PEACH: A sentence-aligned Parallel English-Arabic Corpus for Healthcare

  • 人工对齐5.2万句英阿医疗文本,确保高质量平行数据
  • 涵盖患者手册与教育材料,共约59万英文与57万阿拉伯词语料
  • 适用于医疗翻译、可读性评估及语言模型微调,适合多领域研究者

本文介绍PEACH,一个包含患者信息手册和教育材料的英阿双语平行语料库。该语料库共含51,671对平行句子,总计约590,517个英文词元和567,707个阿拉伯词语元,平均句长为9.52至11.83词。作为人工对齐的语料库,PEACH具有黄金标准价值,可用于对比语言学、翻译研究及自然语言处理。其应用包括构建双语词典、为特定领域机器翻译适配大语言模型、评估医疗机器翻译用户感知、分析患者资料可读性与通俗性,以及作为翻译教学资源。PEACH已公开可获取。

原文摘要 · Abstract (English)

This paper introduces PEACH, a sentence-aligned parallel English-Arabic corpus of healthcare texts encompassing patient information leaflets and educational materials. The corpus contains 51,671 parallel sentences, totaling approximately 590,517 English and 567,707 Arabic word tokens. Sentence lengths vary between 9.52 and 11.83 words on average. As a manually aligned corpus, PEACH is a gold-standard corpus, aiding researchers in contrastive linguistics, translation studies, and natural language processing. It can be used to derive bilingual lexicons, adapt large language models for domain-specific machine translation, evaluate user perceptions of machine translation in healthcare, assess patient information leaflets and educational materials' readability and lay-friendliness, and as an educational resource in translation studies. PEACH is publicly accessible.

医疗翻译双语语料英阿对照语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。