用大模型和自建数据集提升爱沙尼亚语文本简化效果
Improving Estonian Text Simplification through Pretrained Language Models and Custom Datasets
- 用GPT-4生成+人工翻译构建爱沙尼亚语简化数据集
- 微调LLaMA在可读性、语法正确性和语义保留上优于NMT模型
- 适合低资源语言文本简化研究者参考复现
本文提出一种基于神经机器翻译(NMT)模型和微调大型语言模型(LLaMA)的文本简化方法。由于爱沙尼亚语现有资源稀缺,研究者通过结合人工翻译语料与GPT-4.0生成的简化文本,构建了一个新数据集。选用OpenNMT作为代表性NMT系统,将LLaMA在该数据集上进行微调。评估结果显示,LLaMA在语法正确性、可读性和语义保留方面均优于OpenNMT。这些结果表明,在低资源语言场景下,大语言模型对文本简化具有显著优势。完整数据集、微调脚本和评估流程已公开发布,以支持可复现性和向其他语言迁移。
原文摘要 · Abstract (English)
This paper presents a method for text simplification based on two neural architectures: a neural machine translation (NMT) model and a fine-tuned large language model (LLaMA). Given the scarcity of existing resources for Estonian, a new dataset was created by combining manually translated corpora with GPT-4.0-generated simplifications. OpenNMT was selected as a representative NMT-based system, while LLaMA was fine-tuned on the constructed dataset. Evaluation shows LLaMA outperforms OpenNMT in grammaticality, readability, and meaning preservation. These results underscore the effectiveness of large language models for text simplification in low-resource language settings. The complete dataset, fine-tuning scripts, and evaluation pipeline are provided in a publicly accessible supplementary package to support reproducibility and adaptation to other languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。