arXiv:2605.04171cs.CL2026-05中稿 · and presented pape…

分析大模型在学术写作中的幻觉问题,发现不同模型表现差异与任务类型密切相关。

Not All That Is Fluent Is Factual: Investigating Hallucinations of Large Language Models in Academic Writing

  • 设计80个学术写作任务,评估四大模型的幻觉程度。
  • Grok/Copilot参考生成强但抽象表达幻觉率高(HI 0.67/0.70)。
  • Gemini/ChatGPT风格控制好,但事实性任务幻觉风险更高(HI 0.53/0.57)

大型语言模型在学术写作中表现出卓越能力,但仍存在幻觉问题。本文针对ChatGPT、Grok、Gemini和Copilot四款主流模型,设计了涵盖参考文献生成、事实解释、摘要生成和写作优化四个类别的80个提示。采用0-5分评分体系评估事实准确性、参考文献有效性、连贯性、风格一致性及学术语气。引入新的加权指标Hallucination Index(HI)量化幻觉程度。研究发现:Grok与Copilot在参考文献生成任务中表现更优,但在抽象或风格化提示下幻觉严重(HI分别为0.67和0.70);Gemini与ChatGPT在语气控制方面较强,但在事实性任务中幻觉风险更高(HI分别为0.53和0.57)。结果表明,幻觉行为不仅受模型架构影响,更取决于任务类型和提示条件。本研究为未来相关研究提供了新方向。

原文摘要 · Abstract (English)

Large Language models (LLMs) show extraordinary abilities, but they are still prone to hallucinations, especially when we use them for generating Academic content. We have investigated four popular LLMs, ChatGPT, Grok, Gemini, and Copilot for hallucinations specifically for academic writing. We have designed 80 prompts across four categories, namely, reference generation, factual explanation, abstract generation, and writing improvement. We evaluated the model using a 0-5 rubric score, which checks factual accuracy, reference validity, coherence, style consistency, and academic tone. A novel weighted metric, Hallucination Index (HI), was introduced to measure hallucination in the responses generated by the models. Some of the most widely used evaluation metrics often fail to check errors which alter sentiment in machine-translated text. We found that Grok and Copilot perform better on reference generation tasks, but they often struggle with abstract or stylistic prompts, with HI values of 0.67 and 0.70, respectively. Whereas, Gemini and ChatGPT have done well with having stronger tone control, but they lack in writing factual tasks and higher hallucination risk with HI scores of 0.53 and 0.57, respectively. Our study found that hallucination behavior does not depend solely on model architecture but also on the type of task and the prompting conditions we are providing. We propose that our work opens new research dimensions for future researchers.

大模型幻觉检测学术写作评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。