用大模型自动生成科研兴趣画像,比人工写更高效。
Scalable Scientific Interest Profiling Using Large Language Models
- 用LLM分析医学主题词(MeSH)或论文摘要生成兴趣画像。
- 机器生成的画像语义相似度中等(BERTScore F1约0.55),但关键词差异大。
- 基于MeSH的画像可读性更好,适合快速了解研究方向。
科研画像有助于发现人才与推动合作,但常滞后更新。亟需自动化、可扩展的方法保持其时效性。本文设计并评估了两种基于大语言模型(LLMs)的科学兴趣画像生成方法:一种基于PubMed摘要,另一种基于医学主题词(MeSH)。研究收集了哥伦比亚大学欧文医学中心595名教师的论文标题、MeSH术语和摘要,获取了167份人工撰写的画像。使用GPT-4o-mini对每位研究人员的兴趣进行总结。通过人工与自动评估,比较机器生成与自我撰写画像的相似性。结果显示,ROUGE-L、BLEU和METEOR得分较低,表明术语重叠少;但BERTScore显示中等语义相似性(MeSH-based为0.542,abstract-based为0.555)。经改写后,相似度提升至F1=0.851。TF-IDF的KL散度分别为8.56(MeSH)和8.58(摘要),说明机器使用不同关键词。人工评审中,77.78%的MeSH画像被评为“良好”或“优秀”,93.44%可读性高,但粒度与准确性不一。专家更偏好67.86%的MeSH衍生画像。大模型在规模化生成科研画像方面具有潜力,其中基于MeSH的方案更具可读性。
原文摘要 · Abstract (English)
Research profiles highlight scientists' research focus, enabling talent discovery and collaborations, but are often outdated. Automated, scalable methods are urgently needed to keep profiles current. We design and evaluate two Large Language Models (LLMs)-based methods to generate scientific interest profiles--one summarizing PubMed abstracts and the other using Medical Subject Headings (MeSH) terms--comparing them with researchers' self-summarized interests. We collected titles, MeSH terms, and abstracts of PubMed publications for 595 faculty at Columbia University Irving Medical Center, obtaining human-written profiles for 167. GPT-4o-mini was prompted to summarize each researcher's interests. Manual and automated evaluations characterized similarities between machine-generated and self-written profiles. The similarity study showed low ROUGE-L, BLEU, and METEOR scores, reflecting little terminological overlap. BERTScore analysis revealed moderate semantic similarity (F1: 0.542 for MeSH-based, 0.555 for abstract-based), despite low lexical overlap. In validation, paraphrased summaries achieved a higher F1 of 0.851. Comparing original and manually paraphrased summaries indicated limitations of such metrics. Kullback-Leibler (KL) Divergence of TF-IDF values (8.56 for MeSH-based, 8.58 for abstract-based) suggests machine summaries employ different keywords than human-written ones. Manual reviews showed 77.78% rated MeSH-based profiling "good" or "excellent," with readability rated favorably in 93.44% of cases, though granularity and accuracy varied. Panel reviews favored 67.86% of MeSH-derived profiles over abstract-derived ones. LLMs promise to automate scientific interest profiling at scale. MeSH-derived profiles have better readability than abstract-derived ones. Machine-generated summaries differ from human-written ones in concept choice, with the latter initiating more novel ideas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。