arXiv:2411.02943cs.CL2024-11被引 13

用大模型分析科研文献,追踪全球可持续发展目标的讨论变化。

Capturing research literature attitude towards Sustainable Development Goals: an LLM-based topic modeling approach

  • 用LLM生成科学摘要嵌入,结合优化器自动建模海量文献主题。
  • 发现数以百计的主题,覆盖2006-2023年科研文献中对可持续目标的态度演变。
  • 结果可交互查看,适用于关注可持续发展研究趋势的研究者。

全球面临多重挑战,威胁人类文明与地球福祉。联合国于2015年提出可持续发展目标(SDGs),旨在2030年前解决这些全球性问题。自然语言处理技术有助于揭示科研文献中关于SDGs的讨论。本文提出一个完全自动化流程:1)从Scopus数据库获取内容,构建涵盖五类SDGs的专用数据集;2)进行主题建模,识别大规模文本集合中的主题;3)通过关键词搜索与主题频次时间序列提取实现主题探索。主题建模中,我们采用扩展至大规模文本语料的BERTopic框架,引入两种创新:一是基于LLM的嵌入计算方法,将科学摘要映射到连续空间;二是超参数优化器,高效为新大数据集找到最优配置。此外,我们构建了交互式仪表盘可视化结果,展示主题随时间的演化。所有结果均可追溯、可探索,提升了主题建模的可解释性。该基于LLM的主题建模流程可捕捉2006–2023年间科研抽象中对可持续发展目标态度的演变。整个工作流具备可复现性,且可推广至任意时间点和任意大规模文本语料。

原文摘要 · Abstract (English)

The world is facing a multitude of challenges that hinder the development of human civilization and the well-being of humanity on the planet. The Sustainable Development Goals (SDGs) were formulated by the United Nations in 2015 to address these global challenges by 2030. Natural language processing techniques can help uncover discussions on SDGs within research literature. We propose a completely automated pipeline to 1) fetch content from the Scopus database and prepare datasets dedicated to five groups of SDGs; 2) perform topic modeling, a statistical technique used to identify topics in large collections of textual data; and 3) enable topic exploration through keywords-based search and topic frequency time series extraction. For topic modeling, we leverage the stack of BERTopic scaled up to be applied on large corpora of textual documents (we find hundreds of topics on hundreds of thousands of documents), introducing i) a novel LLM-based embeddings computation for representing scientific abstracts in the continuous space and ii) a hyperparameter optimizer to efficiently find the best configuration for any new big datasets. We additionally produce the visualization of results on interactive dashboards reporting topics' temporal evolution. Results are made inspectable and explorable, contributing to the interpretability of the topic modeling process. Our proposed LLM-based topic modeling pipeline for big-text datasets allows users to capture insights on the evolution of the attitude toward SDGs within scientific abstracts in the 2006-2023 time span. All the results are reproducible by using our system; the workflow can be generalized to be applied at any point in time to any big corpus of textual documents.

可持续发展主题建模大模型文献分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。