arXiv:2412.18655cs.CL2024-12被引 4

提出兼顾可读性与连贯性的文档级文本简化系统,提升真实场景下的可读性。

Simple is not Enough: Document-level Text Simplification using Readability and Coherence

  • 融合简洁性、可读性和连贯性多目标训练,改进文档级简化效果。
  • 扩展专业标注语料库,构建复杂-简单-可读性标签三元组用于训练。
  • 在零样本、少样本和微调设置下验证模型,适合需要高连贯性输出的场景。

本文提出 SimDoc 系统,一种考虑简洁性、可读性和篇章连贯性的文本简化模型。过去十年的文本简化研究主要聚焦于句子级别,而实际应用中用户更依赖段落或文档级别的简化结果。我们采用专业标注语料库进行初始微调,并在训练中引入多个目标,同时优化简洁性、可读性和连贯性。贡献包括将现有标注关联为(复杂文本, 简化文本, 可读性标签)三元组,以支持可读性建模;在文档级简化语料库上评估模型在零样本、少样本和微调设置下的表现,展示新方法的有效性;并通过详细输出分析,揭示文档级简化的挑战。

原文摘要 · Abstract (English)

In this paper, we present the SimDoc system, a simplification model considering simplicity, readability, and discourse aspects, such as coherence. In the past decade, the progress of the Text Simplification (TS) field has been mostly shown at a sentence level, rather than considering paragraphs or documents, a setting from which most TS audiences would benefit. We propose a simplification system that is initially fine-tuned with professionally created corpora. Further, we include multiple objectives during training, considering simplicity, readability, and coherence altogether. Our contributions include the extension of professionally annotated simplification corpora by the association of existing annotations into (complex text, simple text, readability label) triples to benefit from readability during training. Also, we present a comparative analysis in which we evaluate our proposed models in a zero-shot, few-shot, and fine-tuning setting using document-level TS corpora, demonstrating novel methods for simplification. Finally, we show a detailed analysis of outputs, highlighting the difficulties of simplification at a document level.

文本简化可读性连贯性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。