LLM依赖论文摘要可能高估研究结论,导致误导性输出。
Will AI be overconfident about academic research findings when reliant on abstracts? (v1)
- 用GPT-OSS 120B对比摘要、讨论、结论中结论强度差异。
- 除人文社科外,摘要和结论的结论强度普遍高于讨论部分。
- 提示研究者慎用仅读摘要的LLM做学术知识总结。
大型语言模型(如ChatGPT、DeepSeek、Gemini)在学术知识发现与摘要中日益普及,但可能因幻觉导致误导。当模型仅访问论文摘要而无全文时,问题更严重。本研究将完整论文提交给OpenAI的GPT-OSS 120B,要求其分别评估摘要、讨论、结论中主要研究结论的强度。结果显示,在社会科学与人文学科之外,摘要与结论中的结论强度普遍高于讨论部分,表明仅依赖摘要可能导致模型对研究结果过度自信,并将此错误认知传递给用户。因此,使用仅基于摘要的LLM进行学术知识挖掘需保持警惕。
原文摘要 · Abstract (English)
Large Language Models (LLMs) like ChatGPT, DeepSeek and Gemini seem to be increasingly used for knowledge discovery, information retrieval, and knowledge summaries, including for academic topics. This can result in users being misled, such as due to hallucinations. These problems may be exacerbated for academic knowledge if LLMs base their answers on journal article abstracts when they lack full text access. To test whether the information content of abstracts can be misleading, full text articles were submitted to the GPT-OSS 120B, an LLM from OpenAI, asking it to assess separately the strength the claims for the main result in the abstract, discussion, and conclusion. Outside the social sciences and humanities, claims tended to be stronger in the abstract and conclusions than the discussion, suggesting that relying on the strength of claims in abstracts would be misleading. Thus, if LLMs ingest abstracts but not full texts, there is a risk that they will be overconfident about the findings and pass it on to users in response to relevant prompts. This is another reason to be cautious about using LLMs for academic-related knowledge discovery and summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。