arXiv:2508.17973cs.CL2025-08中稿 · INLG 2025被引 2

首个德语可读性可控段落改写数据集与模型,支持五级难度精准适配。

German4All -- A Dataset and Model for Readability-Controlled Paraphrasing in German

  • 基于GPT-4生成2.5万+条多级可读性对齐改写样本
  • 训练出在德语简化任务中达顶尖性能的开源模型
  • 适合无障碍文本生成、教育内容定制等场景

跨不同复杂度水平的文本改写能力对于生成可访问性文本至关重要,可针对不同读者群体进行定制。为此,我们提出German4All,首个大规模德语段落级可读性控制改写数据集,涵盖五个可读性等级,包含超过25,000个样本。该数据集通过GPT-4自动生成,并经过人工与大模型双重评估验证。基于German4All,我们训练了一个开源的可读性控制改写模型,在德语文本简化任务中达到当前最优表现,实现更精细、读者导向的文本适配。我们开源了数据集和模型,以促进多层级改写研究。

原文摘要 · Abstract (English)

The ability to paraphrase texts across different complexity levels is essential for creating accessible texts that can be tailored toward diverse reader groups. Thus, we introduce German4All, the first large-scale German dataset of aligned readability-controlled, paragraph-level paraphrases. It spans five readability levels and comprises over 25,000 samples. The dataset is automatically synthesized using GPT-4 and rigorously evaluated through both human and LLM-based judgments. Using German4All, we train an open-source, readability-controlled paraphrasing model that achieves state-of-the-art performance in German text simplification, enabling more nuanced and reader-specific adaptations. We opensource both the dataset and the model to encourage further research on multi-level paraphrasing

可读性控制德语处理文本简化数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。