arXiv:2507.03415cs.CL2025-07被引 2

用语义有意义的表示训练自回归模型,生成高质量同义文本。

SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation

  • 以语义表征作为初始嵌入,引导自回归模型生成同义句。
  • 在无监督场景下达到当前最佳效果,优于现有方法。
  • 揭示常用评估指标(如BLEU、BERTScore)可靠性不足。

本文提出语义有意义的因果语言建模(SMCLM),一种自监督方法,用于训练自回归模型生成语义等价文本。该方法在自回归训练与生成过程中使用语义有意义的文本表示作为初始嵌入。大量实验证明,SMCLM使自回归模型具备学习鲁棒且高质量同义句生成的能力。所提方法在无监督场景下表现优异,性能可与有监督方法媲美,并达到当前最优水平。本文还构建了一套全面的自动评估指标,覆盖多种生成同义句的评价维度。同时指出,广泛使用的评估指标如BLEU、ROUGE和BERTScore存在可靠性问题。

原文摘要 · Abstract (English)

This article introduces semantically meaningful causal language modeling (SMCLM), a selfsupervised method of training autoregressive models to generate semantically equivalent text. Our approach involves using semantically meaningful text representation as an initial embedding in the autoregressive training and generation processes. The extensive empirical study demonstrates that the SMCLM approach makes autoregressive models capable of learning robust and high-quality paraphrase generation. The proposed method is competitive with the supervised method and achieves state-of-the-art results in unsupervised approaches. This article also presents a comprehensive set of automatic metrics that cover a wide range of autogenerated paraphrase evaluation aspects. Simultaneously, this article highlights the low reliability of the metrics that are widely used in paraphrase generation evaluation, including BLEU, ROUGE, and BERTScore.

文本生成自回归模型同义改写评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。