首个学术转大众文本改写数据集,让论文更易懂。
VTechAGP: An Academic-to-General-Audience Text Paraphrase Dataset and Benchmark Models
- 构建文档级学术-大众文本改写数据集VTechAGP
- 轻量模型DSPT5在改写任务中表现优于大模型
- 适合科研传播、科普写作与AI可读性研究
现有文本简化或改写数据集多聚焦通用领域的句子级生成,且缺乏领域知识。本文发布首个学术转大众文本改写数据集VTechAGP,包含8个学院、25年以上学位论文与摘要的文档级配对数据。我们提出动态软提示生成模型DSPT5,采用对比生成损失函数学习动态提示中的关键词向量,并在推理时结合语义与结构层面的众包采样策略筛选最优输出。实验表明,当前主流大模型表现不佳,而轻量级的DSPT5可取得竞争力结果。据我们所知,这是首个针对学术转大众文本改写的基准数据集与解决方案。模型将在论文接受后公开。
原文摘要 · Abstract (English)
Existing text simplification or paraphrase datasets mainly focus on sentence-level text generation in a general domain. These datasets are typically developed without using domain knowledge. In this paper, we release a novel dataset, VTechAGP, which is the first academic-to-general-audience text paraphrase dataset consisting of document-level these and dissertation academic and general-audience abstract pairs from 8 colleges authored over 25 years. We also propose a novel dynamic soft prompt generative language model, DSPT5. For training, we leverage a contrastive-generative loss function to learn the keyword vectors in the dynamic prompt. For inference, we adopt a crowd-sampling decoding strategy at both semantic and structural levels to further select the best output candidate. We evaluate DSPT5 and various state-of-the-art large language models (LLMs) from multiple perspectives. Results demonstrate that the SOTA LLMs do not provide satisfactory outcomes, while the lightweight DSPT5 can achieve competitive results. To the best of our knowledge, we are the first to build a benchmark dataset and solutions for academic-to-general-audience text paraphrase dataset. Models will be public after acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。