评测六款商用AI聊天机器人新闻回答能力,发现其准确率受语言和查询质量影响显著。
Evaluating Commercial AI Chatbots as News Intermediaries

- 在多语言新闻事实问答中,采用检索-生成流水线的聊天机器人表现参差。
- 对已发生事件回答准确率达90%以上,但自由作答时下降11%-17%。
- 对含错误前提的问题极易误信,且存在区域偏见与检索依赖问题。
AI聊天机器人正迅速改变人们获取新闻的方式,但此前尚无研究系统评估这些基于专有搜索集成与检索-合成流程的系统,在跨语言、跨区域情境下处理新兴事实的能力。我们对六款聊天机器人(Gemini 3 Flash 和 Pro、Grok 4、Claude 4.5 Sonnet、GPT-5 和 GPT-4o mini)进行了为期14天(2026年2月9日至22日)的评估,涵盖2,100个来自六大地区服务(美加、阿拉伯语、非洲、印地语、俄语、土耳其语)同日BBC新闻的事实性问题。最佳系统在事件发生数小时后的问题上,多项选择准确率超90%;但在自由作答评估中,准确率下降11%-13%,整体下降16%-17%。我们识别出三种失败模式:第一,所有模型在印地语任务上表现最差(79%),低于其他语言(89%-91%),引用显示英语内容存在检索偏见;第二,超过70%的错误源于检索失败,而非推理;第三,面对含细微虚假前提的问题,准确率从88%-96%骤降至19%-70%,最脆弱模型接受虚构事实比例达64%。此外,还发现检测-准确率悖论:最强假前提检测器在对抗性准确率(弃答率)上仅排第二,而较弱检测器排名第一,表明检测与答案恢复是部分独立能力。总体表明,高准确率可能掩盖系统性区域不公、对检索基础设施的高度依赖,以及对真实用户不完美提问的脆弱性。
原文摘要 · Abstract (English)
AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their proprietary search integrations and retrieval-synthesis pipelines, handle emerging facts across languages and regions. We present a 14-day (February 9-22, 2026) evaluation of six AI chatbots (Gemini 3 Flash and Pro, Grok 4, Claude 4.5 Sonnet, GPT-5 and GPT-4o mini) on 2,100 factual questions derived from same-day BBC News reporting across six regional services (US & Canada, Arabic, Afrique, Hindi, Russian, Turkish). The best systems achieve over 90% multiple-choice accuracy on questions about events reported hours earlier. The same systems, however, lose 11-13% under free-response evaluation, and 16-17% across the cohort. We further characterize three failure patterns. First, every model achieves its lowest accuracy on Hindi (79% vs. 89-91% elsewhere) and citations indicate an Anglophone retrieval bias (e.g., models answering Hindi queries cite English Wikipedia more than any Hindi outlet). Second, retrieval, not reasoning, failures drive over 70% of all errors. When models retrieve a correct source, they often extract the correct answer; the problem is to land on the right source in the first place. Third, models achieving 88-96% accuracy on well-formed questions drop to 19-70% when questions contain subtle false premises, with the most vulnerable model accepting fabricated facts 64% of the time. We also identify a detection-accuracy paradox: the best false-premise detector ranks second in adversarial accuracy (abstention rate), while a weaker detector ranks first, showing that premise detection and answer recovery are partially independent capabilities. Overall, these suggest that high accuracy can mask systematic regional inequity, near-total dependence on retrieval infrastructure, and vulnerability to imperfect queries real users pose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。