arXiv:2508.16303cs.CL2025-08被引 7

构建了超大规模日英专利双语语料库,助力机器翻译与专利分析。

JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus

  • 基于专利家族信息,通过双语对齐提取140万文档对与3.5亿句对。
  • 新增3亿句对后,专利翻译准确率提升20个BLEU点。
  • 适合从事法律文本翻译、跨语言专利挖掘的研究者使用。

我们构建了JaParaPat(日英专利申请平行语料库),涵盖2000至2021年间日本和美国发布的专利申请中的超过3亿条日英句子对。数据来源于日本专利局(JPO)和美国专利商标局(USPTO)的未审查专利公开文件,并通过欧洲专利局(EPO)维护的DOCDB文献数据库获取专利家族信息。基于专利家族关系,我们提取了约140万组相互翻译的日英文档对,并采用基于翻译的句子对齐方法,从文档对中抽取约3.5亿句对,初始翻译模型由基于词典的对齐方法引导。实验表明,在原有2200万网络句对基础上增加超过3亿条专利句对后,专利翻译准确率提升20个BLEU点。

原文摘要 · Abstract (English)

We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United States from 2000 to 2021. We obtained the publication of unexamined patent applications from the Japan Patent Office (JPO) and the United States Patent and Trademark Office (USPTO). We also obtained patent family information from the DOCDB, that is a bibliographic database maintained by the European Patent Office (EPO). We extracted approximately 1.4M Japanese-English document pairs, which are translations of each other based on the patent families, and extracted about 350M sentence pairs from the document pairs using a translation-based sentence alignment method whose initial translation model is bootstrapped from a dictionary-based sentence alignment method. We experimentally improved the accuracy of the patent translations by 20 bleu points by adding more than 300M sentence pairs obtained from patent applications to 22M sentence pairs obtained from the web.

专利数据双语语料机器翻译法律NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。