arXiv:2605.23715cs.CL2026-05被引 2

NLG评估从语言学走向机器学习,未来将更关注影响、质量与安全。

NLG Evaluation: Past, Present, Future

  • 从语言学主导转向机器学习驱动的实验评估
  • 近年兴起大模型作为评判者(LLM-as-Judge)的新方法
  • 未来需重视生成内容的影响、质量及安全性

自然语言生成(NLG)的评估自1990年以来发生了巨大变化,并将在未来持续演进。1990年,当NLG与语言学紧密关联时,尚无现代意义上的正式实验评估。到2026年,随着NLG与机器学习深度结合,实验评估已成为研究的基础。这一时期发展出多种评估技术,包括最近出现的以大模型为评判者的(LLM-as-Judge)方法。预计未来NLG评估将继续演进,特别是对生成内容的影响、定性特征和安全性评估将变得日益重要,因为越来越多的人将日常使用NLG技术。

原文摘要 · Abstract (English)

Natural Language Generation (NLG) evaluation has changed dramatically since 1990, and will continue to evolve in the future. In 1990, when NLG had close ties to linguistics, there was very little formal experimental evaluation in the modern sense. In 2026, when NLG is closely linked to machine learning, experimental evaluation is expected and indeed fundamental to research. Many evaluation techniques were developed over this period, including most recently LLM-as-Judge. I expect NLG evaluation will continue to evolve in the future. In particular, impact, qualitative, and safety evaluation will become more important as large numbers of people routinely use NLG technology.

NLG评估大模型评测安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。