构建生物医学领域摘要数据集,评估大模型生成科学极简摘要能力
Not too long do read: Evaluating LLM-generated extreme scientific summaries
- 基于作者自评摘要构建新数据集BiomedTLDR
- 大模型摘要更依赖原文词汇和结构,偏提取式而非抽象式
- 适合关注科学传播与大模型摘要质量的研究者
高质量的科学极简摘要(TLDR)有助于高效传递科研成果。大语言模型在生成此类摘要方面表现如何?与人类专家有何差异?然而,缺乏高质量、全面的科学TLDR数据集,制约了对LLMs摘要能力的开发与评估。为此,我们提出新数据集BiomedTLDR,包含来自科研论文的大量研究者自撰摘要,这些摘要源于作者常在参考文献项旁附带的评论。我们基于论文摘要测试多个开源大模型生成TLDR的能力。分析显示,尽管部分模型可生成类人化摘要,但总体上,大模型更倾向于保留原文词汇选择与修辞结构,相比人类更偏向提取式而非抽象式摘要。代码与数据集已公开于https://github.com/netknowledge/LLM_summarization (Lyu and Ke, 2025)。
原文摘要 · Abstract (English)
High-quality scientific extreme summary (TLDR) facilitates effective science communication. How do large language models (LLMs) perform in generating them? How are LLM-generated summaries different from those written by human experts? However, the lack of a comprehensive, high-quality scientific TLDR dataset hinders both the development and evaluation of LLMs' summarization ability. To address these, we propose a novel dataset, BiomedTLDR, containing a large sample of researcher-authored summaries from scientific papers, which leverages the common practice of including authors' comments alongside bibliography items. We then test popular open-weight LLMs for generating TLDRs based on abstracts. Our analysis reveals that, although some of them successfully produce humanoid summaries, LLMs generally exhibit a greater affinity for the original text's lexical choices and rhetorical structures, hence tend to be more extractive rather than abstractive in general, compared to humans. Our code and datasets are available at https://github.com/netknowledge/LLM_summarization (Lyu and Ke, 2025).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。