构建首个大规模越南语句子改写数据集,助力自然语言处理研究
A Large-Scale Benchmark for Vietnamese Sentence Paraphrases
- 通过自动生成与人工评估结合的方式构建120万对高质量语句改写数据
- 在多个模型上验证效果,包括BART、T5及GPT-4o等大语言模型
- 为越南语自然语言处理提供重要基准,适合语义理解与文本生成研究者
本文提出ViSP,一个面向越南语句子改写的高质量大规模数据集,包含120万条来自不同领域的原始句-改写句对。数据集采用混合方法构建,结合自动改写生成与人工评估,确保数据质量。我们使用回译、EDA等方法以及BART、T5等基线模型,并测试了GPT-4o、Gemini-1.5、Aya、Qwen-2.5和Meta-Llama-3.1等大语言模型。据我们所知,这是首个针对越南语改写的大型研究。希望该数据集与研究成果能为未来越南语改写任务的研究与应用奠定基础。
原文摘要 · Abstract (English)
This paper presents ViSP, a high-quality Vietnamese dataset for sentence paraphrasing, consisting of 1.2M original-paraphrase pairs collected from various domains. The dataset was constructed using a hybrid approach that combines automatic paraphrase generation with manual evaluation to ensure high quality. We conducted experiments using methods such as back-translation, EDA, and baseline models like BART and T5, as well as large language models (LLMs), including GPT-4o, Gemini-1.5, Aya, Qwen-2.5, and Meta-Llama-3.1 variants. To the best of our knowledge, this is the first large-scale study on Vietnamese paraphrasing. We hope that our dataset and findings will serve as a valuable foundation for future research and applications in Vietnamese paraphrase tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。