arXiv:2502.11442cs.IRcs.AI2025-02被引 6

让对话搜索能跨轮次结合图文提问,更准理解复杂需求。

Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding

  • 用多轮图文交互逐步澄清用户意图,支持复杂查询
  • 在13000+多轮对话上训练,提升MRR达12.88%
  • 适合需要精细交互的智能搜索与对话系统

对话式查询澄清通过互动对话帮助用户优化搜索请求,提升检索效果。传统方法依赖文本澄清问题,难以捕捉涉及视觉属性的复杂偏好。尽管近期研究已探索图文单轮澄清,但无法充分支持用户意图在多轮中的渐进式细化。为此,我们提出多轮多模态澄清问题(MMCQ)任务,融合文本与视觉模态,在多轮对话中优化用户查询。为支持该任务,我们构建了大规模数据集ClariMM,包含超过13,000个多轮交互和33,000个问答对,涵盖多模态澄清问题。我们提出Mario框架,采用两阶段排序策略:先用BM25进行初步检索,再通过整合对话历史中文本与视觉信息的多模态生成重排模型进行精排。实验表明,多轮多模态澄清显著优于单模态及单轮方法,使MRR提升12.88%,尤其在长对话中优势明显,验证了渐进式细化对复杂查询的价值。

原文摘要 · Abstract (English)

Conversational query clarification enables users to refine their search queries through interactive dialogue, improving search effectiveness. Traditional approaches rely on text-based clarifying questions, which often fail to capture complex user preferences, particularly those involving visual attributes. While recent work has explored single-turn multi-modal clarification with images alongside text, such methods do not fully support the progressive nature of user intent refinement over multiple turns. Motivated by this, we introduce the Multi-turn Multi-modal Clarifying Questions (MMCQ) task, which combines text and visual modalities to refine user queries in a multi-turn conversation. To facilitate this task, we create a large-scale dataset named ClariMM comprising over 13k multi-turn interactions and 33k question-answer pairs containing multi-modal clarifying questions. We propose Mario, a retrieval framework that employs a two-phase ranking strategy: initial retrieval with BM25, followed by a multi-modal generative re-ranking model that integrates textual and visual information from conversational history. Our experiments show that multi-turn multi-modal clarification outperforms uni-modal and single-turn approaches, improving MRR by 12.88%. The gains are most significant in longer interactions, demonstrating the value of progressive refinement for complex queries.

多轮对话图文交互查询澄清多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。