对比中文事实搜索与AI回答的可靠性差异,发现答得少的模型反而更准。
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
- 基于真实搜索日志构建中文事实问答数据集,评估九个系统表现
- 搜索引擎答肯定的准确率73.2%~78.9%,但答“是”的比例超83%
- 所有系统对“否”类问题表现更差,且高关注度地区易受误导
搜索引擎和基于AI的系统日益成为获取事实信息的主要途径,但其在真实信息检索场景下的可靠性仍难评估。本文针对中文网络生态,基于真实中文搜索日志构建查询式事实核查数据集,对比传统搜索引擎、独立大语言模型及集成搜索的AI概览系统在中文事实性是非问题上的表现。评估系统在有无证据支持下的正确、错误或不确定判断。结果发现,当系统给出确定答案时,准确率在73.2%至78.9%之间,但回答频率差异显著:搜索引擎在超过83%的查询中给出明确答复,而Qwen-Max仅在不足一半的查询中如此。同时存在一致性的极性偏差:所有系统对“是”类问题的表现优于“否”类。结合百度指数数据,识别出健康相关搜索热度较高的省份,可能预示更高信息误导风险。总体表明,可靠性不仅取决于答对与否,还受回答频率、负向陈述处理能力及信息需求热点影响。
原文摘要 · Abstract (English)
Search engines and AI-powered systems increasingly mediate access to factual information, yet their reliability remains difficult to evaluate in realistic information-seeking settings. We study this problem in the Chinese web ecosystem by constructing a query-based fact-checking dataset from real Chinese search logs and comparing nine systems across traditional search engines, standalone large language models, and search-integrated AI Overviews. Focusing on factual Chinese-language factual Yes/No questions, we evaluate whether systems provide correct, incorrect, or uncertain decisions against evidence-derived ground truth. We find that systems are similarly accurate when they provide definitive answers, but differ sharply in how often they do so. Conditional accuracy ranges from 73.2% to 78.9%, yet search engines answer definitively on over 83% of queries, while Qwen-Max does so on fewer than half. We also find a consistent polarity gap: all systems perform better on yes-labeled queries than on no-labeled queries. We also use Baidu Index data to identify Chinese provinces with higher health-related search attention, which may indicate greater potential exposure to misinformation. Overall, our results show that reliability depends not only on whether systems are correct when they answer, but also on how often they answer, how they handle negative claims, and where information demand may increase exposure risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。