arXiv:2510.22242cs.IRcs.AI2025-10被引 2

测试大模型查论文的靠谱程度,发现它们常答错、乱编。

PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading

  • 构建四类学术任务基准,模拟真实使用场景评测大模型
  • 多参考文献查询错误率高达98%,内容提取失败率超90%
  • 适合关注AI科研助手可靠性的人看

大型语言模型(LLMs)日益充当研究助手,但其在学术任务中的可靠性尚未充分评估。本文提出PaperAsk基准,系统评估GPT-4o、GPT-5和Gemini-2.5-Flash在引用检索、内容提取、论文发现和主张验证四项关键研究任务上的表现。在通过网页界面进行的真实使用条件下,控制实验显示:多参考文献查询中引用检索失败率达48–98%,章节级内容提取失败率为72–91%,主题论文发现的F1分数低于0.32,漏检超过60%相关文献。人工分析表明,失败主因是检索上下文无控扩展及模型更倾向语义相关文本而非任务指令。不同模型表现各异:ChatGPT常选择不回答以避错,Gemini则生成流畅但虚构的内容。为此,我们基于PaperAsk数据训练轻量级可靠性分类器以识别不可靠输出。PaperAsk提供可复现、可诊断的框架,推动基于LLM的学术辅助系统可靠性评估发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) increasingly serve as research assistants, yet their reliability in scholarly tasks remains under-evaluated. In this work, we introduce PaperAsk, a benchmark that systematically evaluates LLMs across four key research tasks: citation retrieval, content extraction, paper discovery, and claim verification. We evaluate GPT-4o, GPT-5, and Gemini-2.5-Flash under realistic usage conditions-via web interfaces where search operations are opaque to the user. Through controlled experiments, we find consistent reliability failures: citation retrieval fails in 48-98% of multi-reference queries, section-specific content extraction fails in 72-91% of cases, and topical paper discovery yields F1 scores below 0.32, missing over 60% of relevant literature. Further human analysis attributes these failures to the uncontrolled expansion of retrieved context and the tendency of LLMs to prioritize semantically relevant text over task instructions. Across basic tasks, the LLMs display distinct failure behaviors: ChatGPT often withholds responses rather than risk errors, whereas Gemini produces fluent but fabricated answers. To address these issues, we develop lightweight reliability classifiers trained on PaperAsk data to identify unreliable outputs. PaperAsk provides a reproducible and diagnostic framework for advancing the reliability evaluation of LLM-based scholarly assistance systems.

大模型评测学术辅助可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。