用户难辨生成回答中的偏见和错误,完整性比正确性更易被察觉。
Can Users Detect Biases or Factual Errors in Generated Responses in Conversational Information-Seeking?
- 通过众包实验对比不同回应变体,评估用户识别问题的能力。
- 用户更易发现回答不完整,而非无法回答问题或事实错误。
- 响应多样性比准确性更能提升用户满意度,提示未来研究重点。
信息查询对话涵盖从简单事实到复杂多面问题的广泛范围。在不熟悉领域进行探索性搜索时,用户因缺乏背景知识而难以验证系统提供的信息,易受误导。本研究探讨对话式信息查询系统生成响应的局限性,揭示其中潜在的不准确、陷阱及偏见。研究关注问题可回答性与响应不完整性的挑战。通过两项众包任务,评估不同系统回应变体对用户体验的影响,重点考察用户识别偏见、错误或不完整回应的能力。分析表明,用户更容易察觉响应不完整,而非问题不可回答;用户满意度主要与响应多样性相关,而非事实正确性。
原文摘要 · Abstract (English)
Information-seeking dialogues span a wide range of questions, from simple factoid to complex queries that require exploring multiple facets and viewpoints. When performing exploratory searches in unfamiliar domains, users may lack background knowledge and struggle to verify the system-provided information, making them vulnerable to misinformation. We investigate the limitations of response generation in conversational information-seeking systems, highlighting potential inaccuracies, pitfalls, and biases in the responses. The study addresses the problem of query answerability and the challenge of response incompleteness. Our user studies explore how these issues impact user experience, focusing on users' ability to identify biased, incorrect, or incomplete responses. We design two crowdsourcing tasks to assess user experience with different system response variants, highlighting critical issues to be addressed in future conversational information-seeking research. Our analysis reveals that it is easier for users to detect response incompleteness than query answerability and user satisfaction is mostly associated with response diversity, not factual correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。