arXiv:2411.11635cs.CLcs.AI2024-11综述被引 3

对比生成式与抽取式摘要技术,评估其在医疗文献知识提取中的应用前景。

Chapter 7 Review of Data-Driven Generative AI Models for Knowledge Extraction from Scientific Literature in Healthcare

  • 对比生成式与抽取式文本摘要方法的原理与效果
  • 基于7项研究分析,发现生成式模型在摘要质量上更具优势
  • 适合关注AI辅助科研文献处理的研究者与医疗信息工作者

本章回顾了从20世纪50年代至今基于自然语言处理的抽象化文本摘要技术的发展历程,重点对比了抽取式摘要与生成式摘要方法。文章梳理了从早期方法到预训练语言模型(如BERT、GPT)的演进过程。通过在PubMed和Web of Science中检索,共识别出60篇相关研究,其中29篇被排除,最终24篇被评估,选取7项研究进行深入分析。章节还包含示例,比较了GPT-3与当前最先进的GPT-4在科学文本摘要任务中的表现。尽管自然语言处理在生成简明文本摘要方面尚未达到完全潜力,但随着对现有问题的逐步解决,这类模型有望在实践中逐步应用。

原文摘要 · Abstract (English)

This review examines the development of abstractive NLP-based text summarization approaches and compares them to existing techniques for extractive summarization. A brief history of text summarization from the 1950s to the introduction of pre-trained language models such as Bidirectional Encoder Representations from Transformer (BERT) and Generative Pre-training Transformers (GPT) are presented. In total, 60 studies were identified in PubMed and Web of Science, of which 29 were excluded and 24 were read and evaluated for eligibility, resulting in the use of seven studies for further analysis. This chapter also includes a section with examples including an example of a comparison between GPT-3 and state-of-the-art GPT-4 solutions in scientific text summarisation. Natural language processing has not yet reached its full potential in the generation of brief textual summaries. As there are acknowledged concerns that must be addressed, we can expect gradual introduction of such models in practise.

文本摘要生成式AI医疗文献NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。