用分阶段方法让大模型更懂长文简化,效果显著提升。
Progressive Document-level Text Simplification via Large Language Models
- 分三步走:先理清篇章结构,再处理话题,最后优化词汇。
- 在多个数据集上超越现有模型,大幅提高简化质量。
- 适合需要高质量长文本简化的研究者和内容创作者。
文本简化研究主要集中在词汇和句子层面,而长文档级简化(DS)仍较少被探索。尽管大语言模型(如ChatGPT)在多项自然语言任务中表现优异,但在DS任务中效果不佳,常将简化误作摘要生成。真正的文档简化需保持全文一致性,并在篇章、句子、词汇层面进行适度简化。人类编辑采用分层简化策略,本文据此提出一种多阶段协作的渐进式简化方法(ProgDS),通过分层分解任务,依次完成篇章级、话题级和词汇级简化。实验表明,ProgDS显著优于现有小模型或直接提示大模型的方法,在文档简化任务上达到新基准。
原文摘要 · Abstract (English)
Research on text simplification has primarily focused on lexical and sentence-level changes. Long document-level simplification (DS) is still relatively unexplored. Large Language Models (LLMs), like ChatGPT, have excelled in many natural language processing tasks. However, their performance on DS tasks is unsatisfactory, as they often treat DS as merely document summarization. For the DS task, the generated long sequences not only must maintain consistency with the original document throughout, but complete moderate simplification operations encompassing discourses, sentences, and word-level simplifications. Human editors employ a hierarchical complexity simplification strategy to simplify documents. This study delves into simulating this strategy through the utilization of a multi-stage collaboration using LLMs. We propose a progressive simplification method (ProgDS) by hierarchically decomposing the task, including the discourse-level, topic-level, and lexical-level simplification. Experimental results demonstrate that ProgDS significantly outperforms existing smaller models or direct prompting with LLMs, advancing the state-of-the-art in the document simplification task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。