arXiv:2410.22349cs.IRcs.AI2024-10被引 28

对比传统搜索,问答引擎常编造信息且引用错误,可信度存疑。

Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses

  • 通过用户测试发现问答引擎有16类缺陷,提出针对性改进方案
  • 自动评估显示三款主流引擎普遍存在幻觉和引用错误问题
  • 开源评测基准,助力透明化评估大模型应用性能

基于大语言模型的问答引擎正取代传统搜索引擎,不仅检索相关来源,还生成带引用的答案摘要。我们对21名用户进行研究,对比问答与传统搜索体验,识别出16项问答引擎缺陷,并提出相应设计建议,关联8项评估指标。通过自动化评估在You.com、Perplexity.ai和BingChat上验证,结果揭示常见问题如频繁幻觉、引用不准确,以及答案置信度差异等特征,与用户研究结论一致。我们发布Answer Engine Evaluation(AEE)评测基准,推动大模型应用的透明化评估。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based applications are graduating from research prototypes to products serving millions of users, influencing how people write and consume information. A prominent example is the appearance of Answer Engines: LLM-based generative search engines supplanting traditional search engines. Answer engines not only retrieve relevant sources to a user query but synthesize answer summaries that cite the sources. To understand these systems' limitations, we first conducted a study with 21 participants, evaluating interactions with answer vs. traditional search engines and identifying 16 answer engine limitations. From these insights, we propose 16 answer engine design recommendations, linked to 8 metrics. An automated evaluation implementing our metrics on three popular engines (You.com, Perplexity.ai, BingChat) quantifies common limitations (e.g., frequent hallucination, inaccurate citation) and unique features (e.g., variation in answer confidence), with results mirroring user study insights. We release our Answer Engine Evaluation benchmark (AEE) to facilitate transparent evaluation of LLM-based applications.

问答引擎幻觉检测评测基准LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。