梳理自然语言生成的核心任务与技术演进。
Natural Language Generation
- 涵盖数据到文本、摘要、图像描述等多类生成任务。
- 指出大模型时代各NLP子领域方法趋同。
- 适合想快速了解NLG全貌的研究者入门。
本文简要概述了自然语言生成(Natural Language Generation, NLG)领域的研究范畴。广义上,NLG指通过自然语言表达信息的系统研究,包括将数据库或知识图谱中的信息转化为文本(数据到文本),以及文本摘要(文本到文本)、图像描述(图像到文本)等任务。作为自然语言处理(NLP)的一个分支,NLG与机器翻译(MT)和对话系统密切相关。部分研究者认为机器翻译不属于NLG,因其不涉及内容选择;而对话系统虽包含NLG,但通常不被归入该范畴,因还包括自然语言理解与对话管理。然而,随着大语言模型(LLMs)的发展,不同NLP子领域在生成自然语言的方法和评估方式上逐渐趋同。
原文摘要 · Abstract (English)
This article provides a brief overview of the field of Natural Language Generation. The term Natural Language Generation (NLG), in its broadest definition, refers to the study of systems that verbalize some form of information through natural language. That information could be stored in a large database or knowledge graph (in data-to-text applications), but NLG researchers may also study summarisation (text-to-text) or image captioning (image-to-text), for example. As a subfield of Natural Language Processing, NLG is closely related to other sub-disciplines such as Machine Translation (MT) and Dialog Systems. Some NLG researchers exclude MT from their definition of the field, since there is no content selection involved where the system has to determine what to say. Conversely, dialog systems do not typically fall under the header of Natural Language Generation since NLG is just one component of dialog systems (the others being Natural Language Understanding and Dialog Management). However, with the rise of Large Language Models (LLMs), different subfields of Natural Language Processing have converged on similar methodologies for the production of natural language and the evaluation of automatically generated text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。