构建首个学术演讲转摘要数据集,助力自动化生成科研报告。
NUTSHELL: A Dataset for Abstract Generation from Scientific Talks
- 收集ACL会议演讲与对应摘要,构建多模态数据集
- 验证新数据集可提升摘要生成质量,显著优于基线模型
- 开源共享,推动语音转摘要技术发展
科学传播正成为自然语言处理的重要方向,尤其在帮助研究人员获取、总结和生成内容方面。其中新兴应用是语音转摘要(SAG),即从录制的学术演讲中自动生成摘要。该技术可帮助研究者高效理解会议演讲,但进展受限于缺乏大规模数据集。为此,本文提出NUTSHELL,一个全新的多模态数据集,包含*ACL*会议演讲及其对应摘要。我们建立了SAG的强基线模型,并通过自动评估指标与人工评价相结合的方式评估生成摘要的质量。结果揭示了当前SAG任务的挑战,同时证明在NUTSHELL上训练能显著提升性能。我们已将NUTSHELL以开放许可(CC-BY 4.0)发布,旨在推动SAG研究发展,促进更优模型与评估方法的诞生。
原文摘要 · Abstract (English)
Scientific communication is receiving increasing attention in natural language processing, especially to help researches access, summarize, and generate content. One emerging application in this area is Speech-to-Abstract Generation (SAG), which aims to automatically generate abstracts from recorded scientific presentations. SAG enables researchers to efficiently engage with conference talks, but progress has been limited by a lack of large-scale datasets. To address this gap, we introduce NUTSHELL, a novel multimodal dataset of *ACL conference talks paired with their corresponding abstracts. We establish strong baselines for SAG and evaluate the quality of generated abstracts using both automatic metrics and human judgments. Our results highlight the challenges of SAG and demonstrate the benefits of training on NUTSHELL. By releasing NUTSHELL under an open license (CC-BY 4.0), we aim to advance research in SAG and foster the development of improved models and evaluation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。