研究图像如何影响对话搜索中问题澄清的效果。
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
- 对比图文与纯文本提问在对话搜索中的表现差异。
- 图像提升问题回答时的参与度,但纯文本更利于准确改写查询。
- 视觉辅助效果因任务类型和用户经验而异,需按场景设计。
对话搜索系统越来越多地使用澄清性问题来优化用户查询并改善搜索体验。以往研究已证明文本型澄清问题能有效提升检索性能与用户体验。尽管图像在多种场景下被证实可提升检索效果,但其在澄清问题中对用户表现的影响仍缺乏探索。本研究通过73名参与者开展用户实验,考察图像在两类搜索任务中的作用:(i) 回答澄清问题,(ii) 查询改写。从多个角度比较了多模态与纯文本澄清问题的表现。结果表明,尽管参与者在回答澄清问题时更偏好图文形式,但在查询改写任务中偏好较为均衡。图像的影响随任务类型和用户经验而变化:在回答问题时,图像有助于维持各水平用户的参与度;在查询改写中,图像促使生成更精确的查询并提升检索性能。有趣的是,在回答澄清问题时,纯文本设置反而表现出更优的用户性能,因其在无图像情况下提供了更全面的文字信息。这些发现为设计高效多模态对话搜索系统提供了重要启示,强调视觉增强的效果具有任务依赖性,应根据具体搜索场景和用户特征进行策略性部署。
原文摘要 · Abstract (English)
Conversational search systems increasingly employ clarifying questions to refine user queries and improve the search experience. Previous studies have demonstrated the usefulness of text-based clarifying questions in enhancing both retrieval performance and user experience. While images have been shown to improve retrieval performance in various contexts, their impact on user performance when incorporated into clarifying questions remains largely unexplored. We conduct a user study with 73 participants to investigate the role of images in conversational search, specifically examining their effects on two search-related tasks: (i) answering clarifying questions and (ii) query reformulation. We compare the effect of multimodal and text-only clarifying questions in both tasks within a conversational search context from various perspectives. Our findings reveal that while participants showed a strong preference for multimodal questions when answering clarifying questions, preferences were more balanced in the query reformulation task. The impact of images varied with both task type and user expertise. In answering clarifying questions, images helped maintain engagement across different expertise levels, while in query reformulation they led to more precise queries and improved retrieval performance. Interestingly, for clarifying question answering, text-only setups demonstrated better user performance as they provided more comprehensive textual information in the absence of images. These results provide valuable insights for designing effective multimodal conversational search systems, highlighting that the benefits of visual augmentation are task-dependent and should be strategically implemented based on the specific search context and user characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。