实验证明,动态搜索能更好帮助用户发现大模型幻觉。
Catch Me if You Search: When Contextual Web Search Results Affect the Detection of Hallucinations
- 让用户自主搜索时,更易识别虚假内容。
- 动态搜索组对真实内容评分更高,自信心更强。
- 适合关注大模型可信度与人机协作的研究者。
随着大语言模型(LLMs)在各类任务中应用增多,其产生错误信息或“幻觉”的问题日益突出,可能带来严重后果。近期将网络搜索结果整合进大模型的做法引发疑问:人们是否利用这些结果来验证生成内容,从而准确识别幻觉?一项在线实验(N=560)研究了提供静态(由模型固定给出)或动态(用户自主搜索)搜索结果,对参与者对模型生成内容(真实、轻微幻觉、严重幻觉)的感知准确性、自我信心以及对模型整体评价的影响,对比无搜索结果的控制组。结果显示,相比控制组,静态与动态条件下的参与者均认为幻觉内容准确性更低,且对模型评价更负面。但动态搜索组在真实内容上的评分更高,整体评估自信心也显著强于静态或控制组。该研究揭示了在实际场景中集成网页搜索功能的实践意义。
原文摘要 · Abstract (English)
While we increasingly rely on large language models (LLMs) for various tasks, these models are known to produce inaccurate content or 'hallucinations' with potentially disastrous consequences. The recent integration of web search results into LLMs prompts the question of whether people utilize them to verify the generated content, thereby accurately detecting hallucinations. An online experiment (N=560) investigated how the provision of search results, either static (i.e., fixed search results provided by LLM) or dynamic (i.e., participant-led searches), affects participants' perceived accuracy of LLM-generated content (i.e., genuine, minor hallucination, major hallucination), self-confidence in accuracy ratings, as well as their overall evaluation of the LLM, as compared to the control condition (i.e., no search results). Results showed that participants in both static and dynamic conditions (vs. control) rated hallucinated content to be less accurate and perceived the LLM more negatively. However, those in the dynamic condition rated genuine content as more accurate and demonstrated greater overall self-confidence in their assessments than those in the static search or control conditions. We highlighted practical implications of incorporating web search functionality into LLMs in real-world contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。