arXiv:2510.13749cs.CL2025-10Conference of the …被引 1

评测聊天助手搜索可信度与回答依据,发现不同模型表现差异大。

Assessing Web Search Credibility and Response Groundedness in Chat Assistants

  • 设计新方法评估助手检索来源可信度和回答依据性
  • Perplexity来源可信度最高,GPT-4o在敏感话题常引用低可信源
  • 为高风险信息场景下的AI系统评估提供基准

聊天助手越来越多地集成网络搜索功能,以获取并引用外部信息。尽管这可能提升答案可靠性,但也增加了放大低可信度信息源的风险。本文提出一种新方法,评估助手在检索过程中的来源可信度及回答与引用来源的契合度。基于5个易产生误导信息的主题,对100个命题进行测试,评估GPT-4o、GPT-5、Perplexity和Qwen Chat的表现。结果显示各助手存在显著差异:Perplexity在来源可信度方面表现最佳,而GPT-4o在敏感话题上更倾向于引用低可信度来源。本研究首次系统比较了主流聊天助手的事实核查行为,为高风险信息环境中评估AI系统提供了基础。

原文摘要 · Abstract (English)

Chat assistants increasingly integrate web search functionality, enabling them to retrieve and cite external sources. While this promises more reliable answers, it also raises the risk of amplifying misinformation from low-credibility sources. In this paper, we introduce a novel methodology for evaluating assistants' web search behavior, focusing on source credibility and the groundedness of responses with respect to cited sources. Using 100 claims across five misinformation-prone topics, we assess GPT-4o, GPT-5, Perplexity, and Qwen Chat. Our findings reveal differences between the assistants, with Perplexity achieving the highest source credibility, whereas GPT-4o exhibits elevated citation of non-credibility sources on sensitive topics. This work provides the first systematic comparison of commonly used chat assistants for fact-checking behavior, offering a foundation for evaluating AI systems in high-stakes information environments.

AI评估可信度信息核查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。