arXiv:2608.29607cs.CVcs.IR2026-08中稿 · EMNLP

首个针对移动端拍照问答的鲁棒性评测基准,揭示真实场景下多模态检索的痛点。

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

论文配图:SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
图 1 · 摘自论文原文
  • 构建53种可控退化条件下的配对数据集,模拟真实手机拍照与输入问题
  • 发现图像模糊严重降低检索性能,而文本错误对联合检索影响有限
  • 提出MOOR方法,通过可靠性感知融合提升噪声环境下的检索鲁棒性

移动端AI作为视觉智能助手,让用户拍摄图片并提问获取信息。拍照问答检索已成为移动端AI最常见入口之一,但图片常模糊,文本问题可能简短或拼写错误。现有基准仅测试干净输入,或未分离拍照问答检索中的配对鲁棒性。因此,我们提出SnapBench,首个针对鲁棒性拍照问答多模态检索的配对基准,涵盖1,145个查询、9,085个候选项,在53种受控退化条件下,附有人工标注。我们评估了16个多模态检索模型,包括双塔编码器和基于嵌入的视觉语言模型。结果表明,图像退化显著降低检索性能,而文本退化主要影响纯文本检索,对联合检索影响较小。干净图像的单模态检索常优于联合检索,说明粗略文本带来拖累,且在噪声输入下缺乏跨模态容错能力。SnapBench为拍照问答场景下的鲁棒检索提供受控测试平台。我们进一步提出MOOR(模态锚定、异常感知、最优重加权),一种简单的自适应融合方法,强调在拍照问答检索中需引入可靠性感知的模态校准。

原文摘要 · Abstract (English)

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, covering dual-tower encoders and embedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack of cross-modal fallback under noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further propose MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.

多模态检索移动端AI鲁棒性评测SnapBench

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。