arXiv:2508.11779cs.CLecon.GN2025-08被引 6

评测大模型处理学术文本能力,发现其在论文评审中表现有限。

A Multi-Task Evaluation of LLMs' Processing of Academic Text Input

  • 设计四项任务评估大模型在学术文本中的角色:复述、比较、评分与反思。
  • 谷歌Gemini虽能可靠概括文本,但评分区分度差,反思缺乏深度。
  • 结果一致表明大模型不宜无限制用于学术同行评审。

大语言模型(LLMs)在科学发现中的作用,尤其是辅助学术同行评审,引发激烈讨论。本文将计算机科学领域常用的研究任务整合为一套有指导、可复现的工作流,评估LLMs对学术文本的处理能力。评估包含四项任务:内容复述、对比、评分与反思,分别要求模型扮演‘信息源’‘裁判’‘知识裁判’和‘合作者’角色,逐步测试其理解科学文本所需智力水平。以顶级信息系统期刊的论文为输入,结合多种文本指标,发现领先模型谷歌Gemini在摘要与改写上表现尚可;通过成对文本比较进行排名时扩展性不足;评分时区分度差;定性反思虽自洽,却难激发有意义研究。该结论在基于语言学、与真实答案对比及人工评估中均一致,且对提示变化稳健。总体不建议未经审查地使用大模型构建同行评审。

原文摘要 · Abstract (English)

How much large language models (LLMs) can aid scientific discovery, notably in assisting academic peer review, is in heated debate. Between a literature digest and a human-comparable research assistant lies their practical application potential. We organize individual tasks that computer science studies employ in separate terms into a guided and robust workflow to evaluate LLMs' processing of academic text input. We employ four tasks in the assessment: content reproduction/comparison/scoring/reflection, each demanding a specific role of the LLM (oracle/judgmental arbiter/knowledgeable arbiter/collaborator) in assisting scholarly works, and altogether testing LLMs with questions that increasingly require intellectual capabilities towards a solid understanding of scientific texts to yield desirable solutions. We exemplify a rigorous performance evaluation with detailed instructions on the prompts. Adopting first-rate Information Systems articles at three top journals as the input texts and an abundant set of text metrics, we record a compromised performance of the leading LLM - Google's Gemini: its summary and paraphrase of academic text is acceptably reliable; using it to rank texts through pairwise text comparison is faintly scalable; asking it to grade academic texts is prone to poor discrimination; its qualitative reflection on the text is self-consistent yet hardly insightful to inspire meaningful research. This evidence against an endorsement of LLMs' text-processing capabilities is consistent across metric-based internal (linguistic assessment), external (comparing to the ground truth), and human evaluation, and is robust to the variations of the prompt. Overall, we do not recommend an unchecked use of LLMs in constructing peer reviews.

大模型评估学术写作同行评审

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。