arXiv:2411.17719cs.CLcs.AI2024-11被引 9

自动将论文转为高质量演示文稿,结构化提取核心内容。

SlideSpawn: An Automatic Slides Generation System for Research Publications

  • 基于论文结构与句子重要性预测生成幻灯片
  • 在650对论文-幻灯片上实现更优生成质量
  • 适合科研汇报、快速理解论文核心

研究论文结构清晰,包含文字、图表、公式、表格等元素,并分为引言、模型、实验等部分,这些特征使其区别于普通文档,有利于高效摘要。本文提出新系统 SlideSpawn,输入论文PDF后自动生成视觉简洁的总结型演示文稿。系统首先将论文PDF转换为带结构信息的XML文档;随后利用在PS5K数据集和自建Aminer 9.5K Insights数据集上训练的机器学习模型,预测每句话的显著性;通过整数线性规划(ILP)选择句子并按语义聚类,每类赋予合适标题;最后将所选句子与其引用的图形元素一同排布生成幻灯片。在650对论文与幻灯片的测试集上,实验表明该系统生成的演示文稿质量更优。

原文摘要 · Abstract (English)

Research papers are well structured documents. They have text, figures, equations, tables etc., to covey their ideas and findings. They are divided into sections like Introduction, Model, Experiments etc., which deal with different aspects of research. Characteristics like these set research papers apart from ordinary documents and allows us to significantly improve their summarization. In this paper, we propose a novel system, SlideSpwan, that takes PDF of a research document as an input and generates a quality presentation providing it's summary in a visual and concise fashion. The system first converts the PDF of the paper to an XML document that has the structural information about various elements. Then a machine learning model, trained on PS5K dataset and Aminer 9.5K Insights dataset (that we introduce), is used to predict salience of each sentence in the paper. Sentences for slides are selected using ILP and clustered based on their similarity with each cluster being given a suitable title. Finally a slide is generated by placing any graphical element referenced in the selected sentences next to them. Experiments on a test set of 650 pairs of papers and slides demonstrate that our system generates presentations with better quality.

论文生成自动化文献摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。