用结构化文献分析提升大模型生成摘要的可靠性
How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description

- 让算法先聚类,再由大模型解读,避免盲目编造
- 大模型直接生成结构时准确率低,依赖人工标注更可信
- 适合科研综述与文献分析场景,提升可解释性
大语言模型(LLMs)可辅助科学文献综述,但常出现虚构参考文献、覆盖不均和主题组织薄弱的问题。本文通过对比六种不同证据与结构水平下的聚类描述生成流程,评估文献计量结构对大模型辅助综述的影响。基于100个已发表的文献计量分析,重建Scopus数据集,提取人工撰写的聚类描述,并从人类一致性、语义覆盖度、聚类质量、图结构质量及参考文献准确性五个维度评估结果。结果显示,大模型生成的内容在语义上接近人工写作,但在缺乏结构引导时不可靠。当文献计量算法预先定义聚类,大模型仅负责解释时,性能显著提升。总体表明,大模型辅助的文献计量分析最适用于混合工作流:算法提供可审计的结构,大模型生成可读性强的描述。
原文摘要 · Abstract (English)
Large language models (LLMs) can support scientific literature synthesis, but remain prone to hallucinated references, uneven coverage, and weakly grounded thematic organization. We evaluate whether bibliometric structure improves LLM-assisted synthesis by comparing six pipelines for generating cluster descriptions under different levels of evidence and structure. Using 100 published bibliometric analyses, we reconstruct Scopus corpora, extract human-written cluster descriptions, and assess outputs by human alignment, semantic coverage, clustering quality, graph quality, and reference grounding. Results show that LLMs produce descriptions semantically close to human-written ones, but are unreliable when asked to infer bibliometric structure from scratch. Performance improves when bibliometric algorithms define the clusters and the LLM interprets them. Overall, LLM-assisted bibliometric synthesis is most promising as a hybrid workflow in which algorithms provide auditable structure and LLMs generate readable descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。