发现Qwen3嵌入模型在对话检索中易受噪声干扰,轻量提示可有效缓解。
Robustness Risk of Conversational Retrieval: Identifying and Mitigating Noise Sensitivity in Qwen3-Embedding Model

- 在短对话式查询下,结构化噪声会显著影响检索排序
- 无提示时,噪声项在前10名结果中占比超40%且语义无关
- 轻量提示可恢复排名稳定性,适合部署场景评估
我们对基于嵌入的对话式检索进行了实证研究,聚焦于短、类对话、弱指定的查询及包含结构化对话痕迹的语料库。以Qwen3-embedding模型为例,发现其存在实际部署中的鲁棒性漏洞:在无查询提示的情况下,结构化的对话噪声虽语义无关,却可能在检索结果中占据主导地位,出现在前10名中占比超过40%。该问题在不同规模模型中持续出现,且在标准干净查询基准上难以察觉,相比早期Qwen版本及其他主流密集检索基线更为严重。进一步表明,轻量级查询提示能显著改变检索行为,有效抑制噪声侵入并恢复排序稳定性。研究揭示了对话检索中一个被忽视的鲁棒性风险,强调需采用反映真实部署复杂性的评估协议。
原文摘要 · Abstract (English)
We present an empirical study of embedding-based retrieval under realistic conversational settings, where queries are short, dialogue-like, and weakly specified, and retrieval corpora contain structured conversational artifacts. Focusing on Qwen3-embedding models, we identify a deployment-relevant robustness vulnerability: under conversational retrieval without query prompting, structured dialogue-style noise can become disproportionately retrievable and intrude into top-ranked results, despite being semantically uninformative. This failure mode emerges consistently across model scales, remains largely invisible under standard clean-query benchmarks, and is significantly more pronounced in Qwen3 than in earlier Qwen variants and other widely used dense retrieval baselines. We further show that lightweight query prompting qualitatively alters retrieval behavior, effectively suppressing noise intrusion and restoring ranking stability. Our findings highlight an underexplored robustness risk in conversational retrieval and underscore the importance of evaluation protocols that reflect the complexities of deployed systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。