研究发现读者难辨论文摘要是否由AI生成,但认为用AI编辑的摘要更可信。
LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- 通过混合方法实验测试读者对AI与人工摘要的辨别能力。
- 用AI编辑的摘要评分高于纯人写或纯AI写的摘要。
- 揭示读者对AI辅助写作的三种不同态度,助力制定学术规范。
大型语言模型(LLMs)越来越多地被用于生成和编辑科学摘要,但其在学术写作中的应用引发了关于信任、质量及披露的疑问。尽管使用日益广泛,人们对读者如何感知由LLM生成的摘要,以及这种感知如何影响对科研成果的评价仍知之甚少。本文通过一项混合方法调查实验,研究了具备机器学习背景的读者能否区分人类与LLM生成的摘要,实际和感知到的LLM参与如何影响对质量和可信度的判断,以及读者对AI辅助写作的态度。结果显示,参与者难以可靠识别出由LLM生成的内容,但对其是否使用了LLM的信念显著影响评价。值得注意的是,经由LLM编辑的摘要得分高于完全由人类撰写或完全由LLM生成的摘要。此外,我们识别出读者对LLM辅助写作的三种不同态度,为科学传播中的披露政策和可接受使用标准提供了依据。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to generate and edit scientific abstracts, yet their integration into academic writing raises questions about trust, quality, and disclosure. Despite growing adoption, little is known about how readers perceive LLM-generated summaries and how these perceptions influence evaluations of scientific work. This paper presents a mixed-methods survey experiment investigating whether readers with ML expertise can distinguish between human- and LLM-generated abstracts, how actual and perceived LLM involvement affects judgments of quality and trustworthiness, and what orientations readers adopt toward AI-assisted writing. Our findings show that participants struggle to reliably identify LLM-generated content, yet their beliefs about LLM involvement significantly shape their evaluations. Notably, abstracts edited by LLMs are rated more favorably than those written solely by humans or LLMs. We also identify three distinct reader orientations toward LLM-assisted writing, offering insights into evolving norms and informing policy around disclosure and acceptable use in scientific communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。