LLM在新闻写作中常过度自信,编造无依据的判断。
Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
- 对比三款模型在真实文档任务中的表现,评估其幻觉率。
- 30%输出含幻觉,Gemini和ChatGPT达40%,远高于NotebookLM的13%。
- 适合新闻从业者警惕AI生成内容的隐性误导,需强化溯源机制。
大型语言模型(LLMs)日益应用于新闻工作流程,但其幻觉倾向威胁新闻的核心实践——来源追溯、归因与准确性。我们在一个包含300份文档的美国TikTok诉讼与政策相关语料库上,评估了ChatGPT、Gemini和NotebookLM三款主流工具在报道型任务中的表现。通过调整提示精确度和上下文长度,并使用分类体系标注句子级输出以衡量幻觉类型与严重程度。结果显示,30%的模型输出至少包含一条幻觉,其中Gemini和ChatGPT的幻觉率约为40%,显著高于NotebookLM的13%。定性分析表明,多数错误并非虚构实体或数字,而是表现为解释性过度自信:模型添加未经支持的对来源的刻画,或将引用观点转化为普遍陈述。这揭示出根本性的认识论错位:新闻要求每项主张都有明确来源,而LLMs即使缺乏证据也生成权威性文本。本文提出适用于新闻领域的幻觉分类扩展,并主张新闻工具应采用强制准确归因的架构,而非仅优化流畅性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in newsroom workflows, but their tendency to hallucinate poses risks to core journalistic practices of sourcing, attribution, and accuracy. We evaluate three widely used tools - ChatGPT, Gemini, and NotebookLM - on a reporting-style task grounded in a 300-document corpus related to TikTok litigation and policy in the U.S. We vary prompt specificity and context size and annotate sentence-level outputs using a taxonomy to measure hallucination type and severity. Across our sample, 30% of model outputs contained at least one hallucination, with rates approximately three times higher for Gemini and ChatGPT (40%) than for NotebookLM (13%). Qualitatively, most errors did not involve invented entities or numbers; instead, we observed interpretive overconfidence - models added unsupported characterizations of sources and transformed attributed opinions into general statements. These patterns reveal a fundamental epistemological mismatch: While journalism requires explicit sourcing for every claim, LLMs generate authoritative-sounding text regardless of evidentiary support. We propose journalism-specific extensions to existing hallucination taxonomies and argue that effective newsroom tools need architectures that enforce accurate attribution rather than optimize for fluency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。