用NMT翻译语法配置,实现多语言文本生成无人工干预。
Evaluation of NMT-Assisted Grammar Transfer for a Multi-Language Configurable Data-to-Text System
- 用NMT+人工初审翻译语法配置,再用于生成
- 在Basketball数据集上生成结果语法正确
- 适合需要多语言批量生成的场景
多语言数据到文本生成的一种方法是提前将源语言的语法配置翻译成目标语言。这些配置随后用于表面实现和文档规划阶段生成输出。本文描述了一种基于规则的NLG实现,其中配置通过神经机器翻译(NMT)结合一次性人工审查进行翻译,并引入跨语言语法依赖模型,构建一个从源数据生成文本的多语言NLG系统,实现在无需人工介入的情况下扩展生成阶段。此外,我们提出一种对自动翻译文本的人工后编辑评估方法。在SportSett:Basketball数据集上的评估表明,该NLG系统表现良好,凸显其在翻译任务中的语法正确性。
原文摘要 · Abstract (English)
One approach for multilingual data-to-text generation is to translate grammatical configurations upfront from the source language into each target language. These configurations are then used by a surface realizer and in document planning stages to generate output. In this paper, we describe a rule-based NLG implementation of this approach where the configuration is translated by Neural Machine Translation (NMT) combined with a one-time human review, and introduce a cross-language grammar dependency model to create a multilingual NLG system that generates text from the source data, scaling the generation phase without a human in the loop. Additionally, we introduce a method for human post-editing evaluation on the automatically translated text. Our evaluation on the SportSett:Basketball dataset shows that our NLG system performs well, underlining its grammatical correctness in translation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。