arXiv:2608.20886cs.CVcs.LG2026-08被引 1

让图像重排序更精准,能理解复杂查询的多维度约束。

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

论文配图:EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
图 1 · 摘自论文原文
  • 将查询解析为六类语义槽的明确证据包,区分必需、禁止和忽略项。
  • 在五个基准上达到顶尖性能,轻量学生模型保留超90%能力。
  • 适合需要精确控制检索结果的场景,如电商或专业搜索。

真实世界图像搜索查询具有多模态和组合性特征,例如‘找这件粉红色的衬衫’需保留特定实体、修改属性并忽略上下文。现有重排序方法要么将多维相关性压缩为不可解释的嵌入,要么依赖易遗漏或虚构细粒度约束的自由形式链式思考。受自然语言处理中评分表与检查清单评估启发,我们将多模态图像重排序重构为语义约束满足问题,提出EviRank:将任意查询(纯文本、纯图像或组合)解析为统一证据包,包含六类语义槽(如实体、属性、关系)的类型化标准,每项标注为必需、禁止或忽略。重排序简化为基于证据的验证,结合确定性评分与证据驱动的列表比较,实现无需训练的统一流程。显式证据还可作为结构化监督,用于可选地蒸馏出轻量级学生模型。在涵盖文本到图像、图像到图像及组合图像检索的五个基准上,EviRank表现达到当前最优,且蒸馏后的学生模型在显著更低成本下仍保持超过90%的教师性能。

原文摘要 · Abstract (English)

Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction problem and propose EviRank, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots (e.g., entities, attributes, relations), each labelled required, forbidden, or ignorable. Re-ranking then reduces to evidence-conditioned verification, combining deterministic rubric scoring and evidence-grounded listwise comparison in a single training-free procedure. The explicit evidence can further serve as structured supervision for optionally distilling a lightweight student. Across five benchmarks spanning text-to-image, image-to-image, and composed image retrieval, EviRank achieves state-of-the-art performance, and the distilled student preserves over 90% of the teacher's capability at substantially lower cost.

图像检索多模态约束推理知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。