arXiv:2502.20582cs.IR2025-02被引 3

构建9万篇顶会论文的结构化摘要数据集,助力科研趋势分析。

CS-PaperSum: A Large-Scale Dataset of AI-Generated Summaries for Scientific Papers

  • 用ChatGPT生成9.2万篇顶会论文的结构化摘要。
  • 关键词重合度与嵌入对齐分析显示关键概念保留良好。
  • 适合研究趋势挖掘、科学发现与智能文献检索的人使用。

计算机科学领域文献快速膨胀,传统数据集仅提供元信息,缺乏捕捉核心贡献与方法的结构化摘要。我们构建了CS-PaperSum,包含来自31个顶级会议的91,919篇论文,通过ChatGPT生成结构化摘要。通过嵌入对齐与关键词重合分析验证摘要质量,结果表明关键概念得以有效保留。案例研究揭示了人工智能研究趋势的变化,包括自监督学习、检索增强生成及多模态AI的兴起。该数据集可支持自动化文献分析、研究趋势预测与人工智能驱动的科学发现,为研究人员、政策制定者及信息检索系统提供重要资源。

原文摘要 · Abstract (English)

The rapid expansion of scientific literature in computer science presents challenges in tracking research trends and extracting key insights. Existing datasets provide metadata but lack structured summaries that capture core contributions and methodologies. We introduce CS-PaperSum, a large-scale dataset of 91,919 papers from 31 top-tier computer science conferences, enriched with AI-generated structured summaries using ChatGPT. To assess summary quality, we conduct embedding alignment analysis and keyword overlap analysis, demonstrating strong preservation of key concepts. We further present a case study on AI research trends, highlighting shifts in methodologies and interdisciplinary crossovers, including the rise of self-supervised learning, retrieval-augmented generation, and multimodal AI. Our dataset enables automated literature analysis, research trend forecasting, and AI-driven scientific discovery, providing a valuable resource for researchers, policymakers, and scientific information retrieval systems.

论文摘要数据集科研趋势AI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。