评估大模型重写论辩文本的语法与语义改进效果
CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Models
- 构建涵盖五十七项指标的四层语言学评估体系
- 模型重写后文本更简洁,平均词长增加且句式合并
- 提升说服力与连贯性,适合论辩生成研究者参考
尽管大语言模型在通用文本生成任务中得到广泛研究,但在文本重写这一相关任务上仍缺乏深入探索,尤其是对模型行为的研究较少。本文聚焦论辩类文本的改写优化,即论辩改进(ArgImp)任务,提出CLEAR评估框架:包含57个映射至词汇、句法、语义和语用四个层次的指标。该框架用于分析多个论辩语料库中大模型重写后的文本质量,并比较不同模型在各语言层面的表现。综合四层分析发现,模型通过缩短文本长度、增加平均词长并合并句子实现论辩改进。总体而言,重写后文本的说服力和连贯性均有所提升。
原文摘要 · Abstract (English)
While LLMs have been extensively studied on general text generation tasks, there is less research on text rewriting, a task related to general text generation, and particularly on the behavior of models on this task. In this paper we analyze what changes LLMs make in a text rewriting setting. We focus specifically on argumentative texts and their improvement, a task named Argument Improvement (ArgImp). We present CLEAR: an evaluation pipeline consisting of 57 metrics mapped to four linguistic levels: lexical, syntactic, semantic and pragmatic. This pipeline is used to examine the qualities of LLM-rewritten arguments on a broad set of argumentation corpora and compare the behavior of different LLMs on this task and analyze the behavior of different LLMs on this task in terms of linguistic levels. By taking all four linguistic levels into consideration, we find that the models perform ArgImp by shortening the texts while simultaneously increasing average word length and merging sentences. Overall we note an increase in the persuasion and coherence dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。